EDBT 2026 Demo / reviewers in the wild / expert
Maya B. Gokhale
dblp:75/273
· DBLP profile ↗
78ranked-venue papers
18as first author
14since 2021 · last 2026
0000-0003-4229-5735ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 70 · 18 first-author · 12 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Theory of computation · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
Ruimin Shi, Maya B. Gokhale, Pei-Hung Lin, Xavier Teruel, Ivy Bo Peng |
Euro-Par (1) | 2 |
| 2026 | Communication Offloading on SmartNIC DPUs: A Quantitative Approach
Jacob Wahlgren, Andong Hu, Roger A. Pearce, Maya B. Gokhale, Ivy Bo Peng |
Euro-Par (1) | 4 |
| 2026 | High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
Ruimin Shi, Gabin Schieffer, Pei-Hung Lin, Maya B. Gokhale, Andreas Herten, Ivy Bo Peng |
IPDPS | 4 |
| 2026 | Optimizing Management of Persistent Data Structures in High-Performance AnalyticsabstractLarge-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application's control, and the significantly high storage footprint of such snapshots. To address these limitations, we presentPrivateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. We integratedPrivateerintoMetall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning.Privateeroptimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system.Privateeralso optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression. Karim Youssef, Keita Iwabuchi, Maya B. Gokhale, Wu-chun Feng, Roger A. Pearce |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
Ruimin Shi, Gabin Schieffer, Maya B. Gokhale, Pei-Hung Lin, Hiren D. Patel, Ivy Bo Peng |
Euro-Par (2) | 3 |
| 2024 | Designing an Energy-Efficient Fully-Asynchronous Deep Learning Convolution EngineabstractIn the face of exponential growth in semiconductor energy usage, there is a significant push towards highly energy-efficient microelectronics design. While the traditional circuit designs typically employ clocks to synchronize the computing operations, these circuits incur significant performance and energy overheads due to their data-independent worst-case operation and complex clock tree networks. In this paper, we explore asynchronous or clockless techniques where clocks are replaced by request, acknowledge handshaking signals. To quantify the potential energy and performance gains of asynchronous logic, we design a highly energy -efficient asynchronous deep learning convolution engine, which uses 87 % of total DL accelerator energy. Our asynchronous design shows 5.06x lower energy and 5.09 x lower delay than the synchronous one. Mattia Vezzoli, Lukas Nel, Kshitij Bhardwaj, Rajit Manohar, Maya B. Gokhale |
DATE | 5 |
| 2024 | Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph ProcessingabstractDisaggregated memory breaks the boundary of monolithic servers to enable memory provisioning on demand. Using network-attached memory to provide memory expansion for memory-intensive applications on compute nodes can improve the overall memory utilization on a cluster and reduce the total cost of ownership. However, current software solutions for leveraging network-attached memory must consume resources on the compute node for memory management tasks. Emerging off-path smartNICs provide general-purpose programmability at low-cost low-power cores. This work provides a general architecture design that enables network-attached memory and offloading tasks onto off-path programmable SmartNIC. We provide a prototype implementation called SODA on Nvidia BlueField DPU. SODA adapts communication paths and data transfer alternatives, pipelines data movement stages, and enables customizable data caching and prefetching optimizations. We evaluate SODA in five representative graph applications on real-world graphs. Our results show that SODA can achieve up to 7.9x speedup compared to node-local SSD and reduce network traffic by 42 % compared to disaggregated memory without SmartNIC offloading at similar or better performance. Jacob Wahlgren, Gabin Schieffer, Maya B. Gokhale, Roger A. Pearce, Ivy Bo Peng |
SBAC-PAD | 3 |
| 2023 | SCCL: An open-source SystemC to RTL translatorabstractWe present SCCL, an open-source tool that translates SystemC designs into synthesizable register-transfer level (RTL). SCCL supports a subset of Accellera's SystemC synthesis standard based on the 2011 revision of C++. We use LLVM's Clang front-end to parse SystemC designs, and a suite of analysis passes to construct a SystemC-specific intermediate abstract syntax tree representation called Hcode. Hcode simplifies translation to other intermediate forms such as FIRRTL as well as direct transcription to SystemVerilog or VHDL. Currently, SCCL provides a translation phase to generate synthesizable SystemVerilog. Distinguishing aspects of SCCL include support for complex templated class descriptions that facilitate concise, parameterized hardware specification; introduction and full support for a new type of synthesizable channel called sc_stream that maps directly to standards such as AXI Stream, and a complete reference implementation targeting the Xilinx Vivado toolchain. We demonstrate SCCL's capabilities with a series of case studies including a highly templated SystemC implementation of the ZFP [1] floating-point codec. All case studies are deployed and executed on a Xilinx Zynq UltraScale+ FPGA platform. Zhuanhao Wu, Maya B. Gokhale, Hiren D. Patel |
FCCM | 2 |
| 2023 | A Quantitative Approach for Adopting Disaggregated Memory in HPC SystemsabstractMemory disaggregation has recently been adopted in data centers to improve resource utilization, motivated by cost and sustainability. Recent studies on large-scale HPC facilities have also highlighted memory underutilization. A promising and non-disruptive option for memory disaggregation is rack-scale memory pooling, where node-local memory is supplemented by shared memory pools. This work outlines the prospects and requirements for adoption and clarifies several misconceptions. We propose a quantitative method for dissecting application requirements on the memory system from the top down in three levels, moving from general, to multi-tier memory systems, and then to memory pooling. We provide a multi-level profiling tool and LBench to facilitate the quantitative approach. We evaluate a set of representative HPC workloads on an emulated platform. Our results show that prefetching activities can significantly influence memory traffic profiles. Interference in memory pooling has varied impacts on applications, depending on their access ratios to memory tiers and arithmetic intensities. Finally, in two case studies, we show the benefits of our findings at the application and system levels, achieving 50% reduction in remote access and 13% speedup in BFS, and reducing performance variation of co-located workloads in interference-aware job scheduling. Jacob Wahlgren, Gabin Schieffer, Maya B. Gokhale, Ivy Bo Peng |
SC | 3 |
| 2022 | Unsupervised Test-Time Adaptation of Deep Neural Networks at the Edge: A Case StudyabstractDeep learning is being increasingly used in mobile and edge autonomous systems. The prediction accuracy of deep neural networks (DNNs), however, can degrade after deployment due to encountering data samples whose distributions are differ-ent than the training samples. To continue to robustly predict, DNNs must be able to adapt themselves post-deployment. Such adaptation at the edge is challenging as new labeled data may not be available, and it has to be performed on a resource-constrained device. This paper performs a case study to evaluate the cost of test-time fully unsupervised adaptation strategies on a real-world edge platform: Nvidia Jetson Xavier NX. In particular, we adapt pretrained state-of-the-art robust DNNs (trained using data augmentation) to improve the accuracy on image classification data that contains various image corruptions. During this prediction-time on-device adaptation, the model parameters of a DNN are updated using a single backpropagation pass while optimizing entropy loss. The effects of following three simple model updates are compared in terms of accuracy, adaptation time and energy: updating only convolutional (Conv-Tune); only fully-connected (FC-Tune); and only batch-norm parameters (BN-Tune). Our study shows that BN-Tune and Conv-Tune are more effective than FC-Tune in terms of improving accuracy for corrupted images data (average of 6.6%, 4.97%, and 4.02%, respectively over no adaptation). However, FC-Tune leads to significantly faster and more energy efficient solution with a small loss in accuracy. Even when using FC-Tune, the extra overheads of on-device fine-tuning are significant to meet tight real-time deadlines (209ms). This study motivates the need for designing hardware-aware robust algorithms for efficient on-device adaptation at the autonomous edge. Kshitij Bhardwaj, James Diffenderfer, Bhavya Kailkhura, Maya B. Gokhale |
DATE | 4 |
| 2022 | ZHW: A Numerical CODEC for Big Data Scientific ComputationabstractDistributed big data in scientific computing presents a major I/O performance bottleneck when exploiting data paral-lelism. Consumer and producer compute nodes are often throttled by saturated data channels when processing large numerical data. We describe ZHW, a hardware implementation of the ZFP numerical CODEC that can greatly reduce I/O pressure caused by large scientific datasets. Our ZHW design overcomes barriers that have prevented prior ZFP-like hardware accelerators from obtaining maximum compression in their implementations. The SystemC ZHW hardware library is available in an open source public repository. We demonstrate the practicality of ZHW by synthesizing our CODEC on an Ultrascale+ FPGA and analyzing performance. Michael Barrow, Zhuanhao Wu, Maya B. Gokhale, Hiren D. Patel, Peter Lindstrom 0001 |
FPT | 4 |
| 2022 | Benchmarking Test-Time Unsupervised Deep Neural Network Adaptation on Edge DevicesabstractThe prediction accuracy of deep neural networks (DNNs) after deployment at the edge can suffer with time due to shifts in the distribution of the new data. To improve robustness of DNNs, they must be able to update themselves. However, DNN adaptation at the edge is challenging due to lack of resources. Recently, lightweight prediction-time unsupervised DNN adaptation techniques have been introduced that improve prediction accuracy of the models for noisy data by re-tuning the batch normalization parameters. This paper performs a comprehensive measurement study of such techniques to quantify their performance and energy on various edge devices as well as find bottlenecks and propose optimization opportunities. Kshitij Bhardwaj, James Diffenderfer, Bhavya Kailkhura, Maya B. Gokhale |
ISPASS | 4 |
| 2022 | Metall: A persistent memory allocator for data-centric analytics
Keita Iwabuchi, Karim Youssef, Kaushik Velusamy, Maya B. Gokhale, Roger A. Pearce |
Parallel Comput. | 4 |
| 2022 | Enabling Scalable and Extensible Memory-Mapped Datastores in UserspaceabstractExascale workloads are expected to incorporate data-intensive processing in close coordination with traditional physics simulations. These emerging scientific, data-analytics and machine learning applications need to access a wide variety of datastores in flat files and structured databases. Programmer productivity is greatly enhanced by mapping datastores into the application process's virtual memory space to provide a unified “in-memory” interface. Currently, memory mapping is provided by system software primarily designed for generality and reliability. However, scalability at high concurrency is a formidable challenge on exascale systems. Also, there is a need for extensibility to support new datastores potentially requiring HPC data transfer services. In this article, we presentUMap, a scalable and extensible userspace service for memory-mapping datastores. Through decoupled queue management, concurrency aware adaptation, and dynamic load balancing,UMapenables application performance to scale even at high concurrency. We evaluateUMapin data-intensive applications, including sorting, graph traversal, database operations, and metagenomic analytics. Our results show thatUMapas a userspace service outperforms an optimized kernel-based service across a wide range of intra-node concurrency by 1.22-1.9${\times}$. We performed two case studies to demonstrateUMap's extensibility. First, a new datastore residing in remote memory is incorporated intoUMapas an application-specific plugin. Second, we present a persistent memory allocatorMetallbuilt atopUMapfor unified storage/memory. Ivy Bo Peng, Maya B. Gokhale, Karim Youssef, Keita Iwabuchi, Roger A. Pearce |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Demystifying the Performance of HPC Scientific Applications on NVM-based Memory SystemsabstractThe emergence of high-density byte-addressable non-volatile memory (NVM) is promising to accelerate data-and compute-intensive applications. Current NVM technologies have lower performance than DRAM and, thus, are often paired with DRAM in a heterogeneous main memory. Recently, byte-addressable NVM hardware becomes available. This work provides a timely evaluation of representative HPC applications from the "Seven Dwarfs" on NVM-based main memory. Our results quantify the effectiveness of DRAM-cached-NVM for accelerating HPC applications and enabling large problems beyond the DRAM capacity. On uncached-NVM, HPC applications exhibit three tiers of performance sensitivity, i.e., insensitive, scaled, and bottlenecked. We identify write throttling and concurrency control as the priorities in optimizing applications. We highlight that concurrency change may have a diverging effect on read and write accesses in applications. Based on these findings, we explore two optimization approaches. First, we provide a prediction model that uses datasets from a small set of configurations to estimate performance at various concurrency and data sizes to avoid exhaustive search in the configuration space. Second, we demonstrate that write-aware data placement on uncached-NVM could achieve 2x performance improvement with a 60% reduction in DRAM usage. Ivy Bo Peng, Kai Wu 0006, Jie Ren 0015, Dong Li 0001, Maya B. Gokhale |
IPDPS | 5 |
| 2020 | On the Memory Underutilization: Exploring Disaggregated Memory on HPC SystemsabstractLarge-scale high-performance computing (HPC) systems consist of massive compute and memory resources tightly coupled in nodes. We perform a large-scale study of memory utilization on four production HPC clusters. Our results show that more than 90% of jobs utilize less than 15% of the node memory capacity, and for 90% of the time, memory utilization is less than 35%. Recently, disaggregated architecture is gaining traction because it can selectively scale up a resource and improve resource utilization. Based on these observations, we explore using disaggregated memory to support memory-intensive applications, while most jobs remain intact on HPC systems with reduced node memory. We designed and developed a user-space remote-memory paging library to enable applications exploring disaggregated memory on existing HPC clusters. We quantified the impact of access patterns and network connectivity in benchmarks. Our case studies of graph-processing and Monte-Carlo applications evaluated the impact of application characteristics and local memory capacity and highlighted the potential of throughput scaling on disaggregated memory. Ivy Bo Peng, Roger A. Pearce, Maya B. Gokhale |
SBAC-PAD | 3 |
| 2019 | Guest Editorial: Special Issue on Reconfigurable Computing and FPGA Technology
René Cumplido, Maya B. Gokhale, Claudia Feregrino-Uribe, Michael Hübner 0001 |
J. Parallel Distributed Comput. | 2 |
| 2018 | Microscope on Memory: MPSoC-Enabled Computer Memory System AssessmentsabstractRecent advances in new memory technologies and packaging options has focused attention on computer memory system design and evaluation. Examples include high bandwidth memories such as Hybrid Memory Cube and HBM, and 3DXPoint non-volatile memory. Emerging memories display a wide range of bandwidths, latencies, and capacities, making it challenging for the computer architect to navigate the design space of potential memory configurations, and for the application developer to assess performance implications of using such memories. In this work, we describe the Logic in Memory Emulator (LiME), a hardware/software tool specially designed for memory system evaluation and experiment. LiME uses the Xilinx Zynq UltraScale+ MPSoC on the ZCU102 board to capture any/all memory access, either from the Processing System (PS) or the Programmable Logic (PL). LiME employs novel loopback circuitry in conjunction with address map relocation to pass memory references from the PS into the PL side. The memory request is looped back into the PS DRAM memory controller and concurrently processed by LiME. We have demonstrated four high value use cases: full external memory access logging and replay, emulation of various memory system latencies by passing the memory request through delay registers before entering the PS memory subsystem, emulation of acceleration engines that can independently access memory, and performance comparison of two CPU architectures with respect to their memory behavior. In this paper we will describe this novel application of state-of-the-art MPSoC embodied by the LiME framework and highlight its uses. Abhishek Kumar Jain, Maya B. Gokhale |
FCCM | 3 |
| 2017 | Argo NodeOS: Toward Unified Resource Management for ExascaleabstractExascale systems are expected to feature hundreds of thousands of compute nodes with hundreds of hardware threads and complex memory hierarchies with a mix of on-package and persistent memory modules. In this context, the Argo project is developing a new operating system for exascale machines. Targeting production workloads using workflows or coupled codes, we improve the Linux kernel on several fronts. We extendthe memory management of Linux to be able to subdivide NUMA memory nodes, allowing better resource partitioning among processes running on the same node. We also add support for memory-mapped access tonode-local, PCIe-attached NVRAM devices and introduce a new scheduling class targeted at parallel runtimes supporting user-level load balancing. These features are unified into compute containers, a containerization approach focused on providing modern HPC applications with dynamic control over a wide range of kernel interfaces. To keep our approach compatible with industrial containerization products, we also identifycontentions points for the adoption of containers in HPC settings. Each NodeOS feature is evaluated by using a set of parallel benchmarks, miniapps, and coupled applications consisting of simulation and data analysis components, running on a modern NUMA platform. We observe out-of-the-box performance improvements easily matching, and often exceeding, those observed with expert-optimized configurations on standard OS kernels. Our lightweight approach to resource management retains the many benefits of a full OS kernel that application programmers have learned to depend on, at the same time providing a set of extensions that can be freely mixed and matched to best benefit particular application components. Swann Perarnau, Judicael A. Zounmevo, Matthieu Dreher, Brian Van Essen, Roberto Gioiosa, Kamil Iskra, Maya B. Gokhale, Kazutomo Yoshii, Pete Beckman |
IPDPS | 7 |
| 2016 | ClearView: Data cleaning for online review miningabstractHow can we automatically clean and curate online reviews to better mine them for knowledge discovery? Typical online reviews are full of noise and abnormalities, hindering semantic analysis and leading to a poor customer experience. Abnormalities include non-standard characters, unstructured punctuation, different/multiple languages, and misspelled words. Worse still, people will leave “junk” text, which is either completely nonsensical, spam, or fraudulent. In this paper, we describe three types of noisy and abnormal reviews, discuss methods to detect and filter them, and, finally, show the effectiveness of our cleaning process by improving the overall distributional characteristics of review datasets. Amanda J. Minnich, Noor Abu-El-Rub, Maya B. Gokhale, Ron Minnich, Abdullah Mueen |
ASONAM | 3 |
| 2016 | RRAM-based TCAMs for pattern searchabstractContent Addressable Memory (CAM) is beneficial to applications that require high-speed pattern searching as it provides fast associative lookup operations. As the amount of data to search continues to grow, reducing power consumption while minimizing the costs for speed and area is the main thread of research in designing large capacity CAMs. In this work, we are presenting an active memory architecture incorporating a searchable resistive memory. The proposed architecture incorporates processing logic in close proximity to the RRAM and the RRAM-based TCAM, where the unit TCAM cell is comprised of five transistors and two memristors. Analyzed and simulated performance (e.g., latency, energy consumption, and storage density) of the RRAM-based TCAM at various technology nodes are presented and compared to those of prior CAM/TCAM designs. Le Zheng, Sangho Shin, Maya B. Gokhale, Kyungmin Kim 0001 |
ISCAS | 4 |
| 2016 | Graph colouring as a challenge problem for dynamic graph processing on distributed systemsabstractAn unprecedented growth in data generation is taking place. Data about larger dynamic systems is being accumulated, capturing finer granularity events, and thus processing requirements are increasingly approaching real-time. To keep up, data-analytics pipelines need to be viable at massive scale, and switch away from static, offline scenarios to support fully online analysis of dynamic systems. This paper uses a challenge problem, graph colouring, to explore massive-scale analytics for dynamic graph processing. We present an event-based infrastructure, and a novel, online, distributed graph colouring algorithm. Our implementation for colouring static graphs, used as a performance baseline, is up to an order of magnitude faster than previous results and handles massive graphs with over 257 billion edges. Our framework supports dynamic graph colouring with performance at large scale better than GraphLab's static analysis. Our experience indicates that online solutions are feasible, and can be more efficient than those based on snapshotting. Scott Sallinen, Keita Iwabuchi, Suraj Poudel, Maya B. Gokhale, Matei Ripeanu, Roger A. Pearce |
SC | 4 |
| 2015 | A Container-Based Approach to OS Specialization for Exascale ComputingabstractFuture exascale systems will impose several conflicting challenges on the operating system (OS) running on the compute nodes of such machines. On the one hand, the targeted extreme scale requires the kind of high resource usage efficiency that is best provided by lightweight OSes. At the same time, substantial changes in hardware are expected for exascale systems. Compute nodes are expected to host a mix of general-purpose and special-purpose processors or accelerators tailored for serial, parallel, compute-intensive, or I/O-intensive workloads. Similarly, the deeper and more complex memory hierarchy will expose multiple coherence domains and NUMA nodes in addition to incorporating nonvolatile RAM. That expected workload and hardware heterogeneity and complexity is not compatible with the simplicity that characterizes high performance lightweight kernels. In this work, we describe the Argo Exascale node OS, which is our approach to providing in a single kernel the required OS environments for the two aforementioned conflicting goals. We resort to multiple OS specializations on top of a single Linux kernel coupled with multiple containers. Judicael A. Zounmevo, Swann Perarnau, Kamil Iskra, Kazutomo Yoshii, Roberto Gioiosa, Brian Van Essen, Maya B. Gokhale, Edgar A. León |
IC2E | 7 |
| 2014 | Faster Parallel Traversal of Scale Free Graphs at Extreme Scale with Vertex DelegatesabstractAt extreme scale, irregularities in the structure of scale-free graphs such as social network graphs limit our ability to analyze these important and growing datasets. A key challenge is the presence of high-degree vertices (hubs), that leads to parallel workload and storage imbalances. The imbalances occur because existing partitioning techniques are not able to effectively partition high-degree vertices. We present techniques to distribute storage, computation, and communication of hubs for extreme scale graphs in distributed memory supercomputers. To balance the hub processing workload, we distribute hub data structures and related computation among a set of delegates. The delegates coordinate using highly optimized, yet portable, asynchronous broadcast and reduction operations. We demonstrate scalability of our new algorithmic technique using Breadth-First Search (BFS), Single Source Shortest Path (SSSP), K-Core Decomposition, and Page-Rank on synthetically generated scale-free graphs. Our results show excellent scalability on large scale-free graphs up to 131K cores of the IBM BG/P, and outperform the best known Graph500 performance on BG/P Intrepid by 15%. Roger A. Pearce, Maya B. Gokhale, Nancy M. Amato |
SC | 2 |
| 2013 | Minerva: Accelerating Data Analysis in Next-Generation SSDsabstractEmerging non-volatile memory (NVM) technologies have DRAM-like latency with storage-like density, offering unique capability to analyze large data sets significantly faster than flash or disk storage. However, the hybrid nature of these NVM technologies such as phase-change memory (PCM) make it difficult to use them to best advantage in the memory-storage hierarchy. These NVMs lack the fast write latency required of DRAM and are thus not suitable as DRAM equivalent on the memory bus, yet their low latency even in random access patterns is not easily exploited over an I/O bus. In this work, we describe an FPGA-based system to execute application-specific operations in the NVM controller and evaluate its performance on two microbenchmarks and a keyvalue store. Our system Minerva1extends the conventional solidstate drive (SSD) architecture to offload data or I/O intensive application code to the SSD to exploit the low latency and high internal bandwidth of NVMs. Performing computation in the FPGA-based NVM storage controller significantly reduces data traffic between the host and storage and serves as an offload engine for data analysis workloads. A runtime library enables the programmer to offload computations to the SSD without dealing with the complications of the underlying architecture and inter-controller communication management. We have implemented a prototype of Minerva on the BEE3 FPGA system. We compare the performance of Minerva to a state of the art PCIe-attached PCM-based SSD. Minerva improves performance by an order of magnitude on two microbenchmarks. Minerva based key-value store performs up to 5.2 M get operations/s and 4.0 M set operations/s which is 7.45× and 9.85× higher than the PCM-based SSD that uses the conventional I/O architecture. This huge improvement comes from the reduction of data transfer between the storage to the host and the FPGA-based data processing in the SSD. Arup De, Maya B. Gokhale, Rajesh K. Gupta 0001, Steven Swanson |
FCCM | 2 |
| 2013 | Scaling Techniques for Massive Scale-Free Graphs in Distributed (External) MemoryabstractWe present techniques to process large scale-free graphs in distributed memory. Our aim is to scale to trillions of edges, and our research is targeted at leadership class supercomputers and clusters with local non-volatile memory, e.g., NAND Flash. We apply an edge list partitioning technique, designed to accommodate high-degree vertices (hubs) that create scaling challenges when processing scale-free graphs. In addition to partitioning hubs, we use ghost vertices to represent the hubs to reduce communication hotspots. We present a scaling study with three important graph algorithms: Breadth-First Search (BFS), K-Core decomposition, and Triangle Counting. We also demonstrate scalability on BG/P Intrepid by comparing to best known Graph500 results [1]. We show results on two clusters with local NVRAM storage that are capable of traversing trillion-edge scale-free graphs. By leveraging node-local NAND Flash, our approach can process thirty-two times larger datasets with only a 39% performance degradation in Traversed Edges Per Second (TEPS). Roger A. Pearce, Maya B. Gokhale, Nancy M. Amato |
IPDPS | 2 |
| 2013 | Scalable metagenomic taxonomy classification using a reference genome databaseabstractMOTIVATION: Deep metagenomic sequencing of biological samples has the potential to recover otherwise difficult-to-detect microorganisms and accurately characterize biological samples with limited prior knowledge of sample contents. Existing metagenomic taxonomic classification algorithms, however, do not scale well to analyze large metagenomic datasets, and balancing classification accuracy with computational efficiency presents a fundamental challenge. RESULTS: A method is presented to shift computational costs to an off-line computation by creating a taxonomy/genome index that supports scalable metagenomic classification. Scalable performance is demonstrated on real and simulated data to show accurate classification in the presence of novel organisms on samples that include viruses, prokaryotes, fungi and protists. Taxonomic classification of the previously published 150 giga-base Tyrolean Iceman dataset was found to take <20 h on a single node 40 core large memory machine and provide new insights on the metagenomic contents of the sample. AVAILABILITY: Software was implemented in C++ and is freely available at http://sourceforge.net/projects/lmat CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sasha Ames, David Hysom, Shea N. Gardner, Maya B. Gokhale, Jonathan E. Allen |
Bioinform. | 5 |
| 2012 | Accelerating a Random Forest Classifier: Multi-Core, GP-GPU, or FPGA?abstractRandom forest classification is a well known machine learning technique that generates classifiers in the form of an ensemble ("forest") of decision trees. The classification of an input sample is determined by the majority classification by the ensemble. Traditional random forest classifiers can be highly effective, but classification using a random forest is memory bound and not typically suitable for acceleration using FPGAs or GP-GPUs due to the need to traverse large, possibly irregular decision trees. Recent work at Lawrence Livermore National Laboratory has developed several variants of random forest classifiers, including the Compact Random Forest (CRF), that can generate decision trees more suitable for acceleration than traditional decision trees. Our paper compares and contrasts the effectiveness of FPGAs, GP-GPUs, and multi-core CPUs for accelerating classification using models generated by compact random forest machine learning classifiers. Taking advantage of training algorithms that can produce compact random forests composed of many, small trees rather than fewer, deep trees, we are able to regularize the forest such that the classification of any sample takes a deterministic amount of time. This optimization then allows us to execute the classifier in a pipelined or single-instruction multiple thread (SIMT) fashion. We show that FPGAs provide the highest performance solution, but require a multi-chip / multi-board system to execute even modest sized forests. GP-GPUs offer a more flexible solution with reasonably high performance that scales with forest size. Finally, multi-threading via Open MP on a shared memory system was the simplest solution and provided near linear performance that scaled with core count, but was still significantly slower than the GP-GPU and FPGA. Brian Van Essen, Chris Macaraeg, Maya B. Gokhale, Ryan Prenger |
FCCM | 3 |
| 2012 | On the Role of NVRAM in Data-intensive Architectures: An EvaluationabstractData-intensive applications are best suited to high-performance computing architectures that contain large quantities of main memory. Creating these systems with DRAM-based main memory remains costly and power-intensive. Due to improvements in density and cost, non-volatile random access memories (NVRAM) have emerged as compelling storage technologies to augment traditional DRAM. This work explores the potential of future NVRAM technologies to store program state at performance comparable to DRAM. We have developed the PerMA NVRAM simulator that allows us to explore applications with working sets ranging up to hundreds of gigabytes per node. The simulator is implemented as a Linux device driver that allows application execution at native speeds. Using the simulator we show the impact of future technology generations of I/O-bus-attached NVRAM on an unstructured-access, level-asynchronous, Breadth-First Search (BFS) graph traversal algorithm. Our simulations show that within a couple of technology generations, a system architecture with local high performance NVRAM will be able to effectively augment DRAM to support highly concurrent data-intensive applications with large memory footprints. However, improvements will be needed in the I/O stack to deliver this performance to applications. The simulator shows that future technology generations of NVRAM in conjunction with an improved I/O runtime will enable parallel data-intensive applications to offload in-memory data structures to NVRAM with minimal performance loss. Brian Van Essen, Roger A. Pearce, Sasha Ames, Maya B. Gokhale |
IPDPS | 4 |
| 2011 | QMDS: A File System Metadata Management Service Supporting a Graph Data Model-Based Query LanguageabstractFile system metadata management has become a bottleneck for many data-intensive applications that rely on high-performance file systems. Part of the bottleneck is due to the limitations of an almost 50 year old interface standard with metadata abstractions that were designed at a time when high-end file systems managed less than 100 MB. Today's high-performance file systems store 7 to 9 orders of magnitude more data, resulting in numbers of data items for which these metadata abstractions are inadequate, such as directory hierarchies unable to handle complex relationships among data. Users of file systems have attempted to work around these inadequacies by moving application-specific metadata management to relational databases to make metadata searchable. Splitting file system metadata management into two separate systems introduces inefficiencies and systems management problems. To address this problem, we propose QMDS: a file system metadata management service that integrates all file system metadata and uses a graph data model with attributes on nodes and edges. Our service uses a query language interface for file identification and attribute retrieval. We present our metadata management service design and architecture and study its performance using a text analysis benchmark application. Results from our QMDS prototype show the effectiveness of this approach. Compared to the use of a file system and relational database, the QMDS prototype shows superior performance for both ingest and query workloads. Sasha Ames, Maya B. Gokhale, Carlos Maltzahn |
NAS | 2 |
| 2011 | Massively parallel acceleration of a document-similarity classifier to detect web attacks
Craig D. Ulmer, Maya B. Gokhale, Brian Gallagher, Philip Top, Tina Eliassi-Rad |
J. Parallel Distributed Comput. | 2 |
| 2010 | Real-Time Classification of Multimedia Traffic Using FPGAabstractReal-time classification of Internet traffic according to application types is vital for network management and surveillance. Identifying emerging applications based on well-known port numbers is no longer reliable. While deep packet inspection (DPI) solutions can be accurate, they require constant updates of signatures and become infeasible for encrypted payload especially in multimedia applications (e.g. Skype). Statistical approaches based on machine learning have thus been considered more promising and robust to encryption, privacy, protocol obfuscation, etc. However, the computation complexity of traffic classification using those statistical solutions is high, which prevents them being deployed in systems that need to manage Internet traffic in real time. This paper proposes a FPGA-based parallel architecture to accelerate the statistical identification of multimedia applications while maintaining high classification accuracy. Specifically, we base our design on the k-Nearest Neighbors (k-NN) algorithm which has been shown to be one of the most accurate machine learning algorithms for Internet traffic classification. To enable high-rate data streaming for real-time classification, we adopt the locality sensitive hashing (LSH) for approximate k-NN. The LSH scheme is carefully designed to achieve high accuracy while being efficient for implementation on FPGA. Processing components in the architecture are optimized to realize high throughput. Extensive experiments and FPGA implementation results show that our design can achieve high accuracy above 99% for classifying three main categories of multimedia applications from Internet traffic while sustaining 80 Gbps throughput for minimum size (40 bytes) packets. Weirong Jiang, Maya B. Gokhale |
FPL | 2 |
| 2010 | FPGA Based Network Traffic Analysis Using Traffic Dispersion PatternsabstractThe problem of Network Traffic Classification (NTC) has attracted significant amount of interest in the research community, offering a wide range of solutions at various levels. The core challenge is in addressing high amounts of traffic diversity found in today's networks. The problem becomes more challenging if a quick detection is required as in the case of identifying malicious network behavior or new applications like peer-to-peer traffic that have potential to quickly throttle the network bandwidth or cause significant damage. Recently, Traffic Dispersion Graphs (TDGs) have been introduced as a viable candidate for NTC. The TDGs work by forming a network wide communication graphs that embed characteristic patterns of underlying network applications. However, these patterns need to be quickly evaluated for mounting real-time response against them. This paper addresses these concerns and presents a novel solution for real-time analysis of Traffic Dispersion Metrics (TDMs) in the TDGs. We evaluate the dispersion metrics of interest and present a dedicated solution on an FPGA for their analysis. We also present analytical measures and empirically evaluate operating effectiveness of our design. The mapped design on Virtex-5 device can process 7.4 million packets/second for a TDG comprising of 10k flows at very high accuracies of over 96%. Maya B. Gokhale, Chen-Nee Chuah |
FPL | 2 |
| 2010 | Multithreaded Asynchronous Graph Traversal for In-Memory and Semi-External MemoryabstractProcessing large graphs is becoming increasingly important for many domains such as social networks, bioinformatics, etc. Unfortunately, many algorithms and implementations do not scale with increasing graph sizes. As a result, researchers have attempted to meet the growing data demands using parallel and external memory techniques. We present a novel asynchronous approach to compute Breadth-First-Search (BFS), Single-Source-Shortest-Paths, and Connected Components for large graphs in shared memory. Our highly parallel asynchronous approach hides data latency due to both poor locality and delays in the underlying graph data storage. We present an experimental study applying our technique to both In-Memory and Semi-External Memory graphs utilizing multi-core processors and solid-state memory devices. Our experiments using synthetic and real-world datasets show that our asynchronous approach is able to overcome data latencies and provide significant speedup over alternative approaches. For example, on billion vertex graphs our asynchronous BFS scales up to 14 x on 16-cores. Roger A. Pearce, Maya B. Gokhale, Nancy M. Amato |
SC | 2 |
| 2009 | Application Experiments: MPPA and FPGAabstractThis paper describes the mapping approach, programmability, and performance of the Ambric massively parallel processor array (MPPA), and compares these aspects to an FPGA. Two application kernels, a trellis decoder, and n-gram frequency counter, were ported to the Ambric development system and an Altera Stratix II. We find that the mapping strategies to Ambric and FPGAs are similar at the high level, but diverge quite a bit in implementation due to differences in granularity between the basic compute units of the two devices. Both require substantial refactoring from the baseline sequential algorithm. The FPGA proved superior in terms of performance, but the Ambric fares significantly better than the FPGA in programmability and ease of application development. Philip Top, Maya B. Gokhale |
FCCM | 2 |
| 2008 | Accelerating Molecular Dynamics Simulations with Reconfigurable ComputersabstractWith advances in reconfigurable hardware, especially field-programmable gate arrays (FPGAs), it has become possible to use reconfigurable hardware to accelerate complex applications such as those in scientific computing. There has been a resulting development of reconfigurable computers, that is, computers that have both general-purpose processors and reconfigurable hardware, as well as memory and high-performance interconnection networks. In this paper, we describe the acceleration of molecular dynamics simulations with reconfigurable computers. We evaluate several design alternatives for the implementation of the application on a reconfigurable computer. We show that a single node accelerated with reconfigurable hardware, utilizing fine-grained parallelism in the reconfigurable hardware design, is able to achieve a speedup of about two times over the corresponding software-only simulation. We then parallelize the application and study the effect of acceleration on performance and scalability. Specifically, we study strong scaling, in which the problem size is fixed. We find that the unaccelerated version actually scales better, because it spends more time in computation than the accelerated version does. However, we also find that a cluster of P accelerated nodes gives better performance than a cluster of 2P unaccelerated nodes. Ronald Scrofano, Maya B. Gokhale, Frans Trouw, Viktor Prasanna 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2008 | A Case Study of Hardware/Software Partitioning of Traffic Simulation on the Cray XD1abstractScientific application kernels mapped to reconfigurable hardware have been reported to have 10times to 100times speedup over equivalent software. These promising results suggest that reconfigurable logic might offer significant speedup on applications in science and engineering. To accurately assess the benefit of hardware acceleration on scientific applications, however, it is necessary to consider the entire application including software components as well as the accelerated kernels. Aspects to be considered include alternative methods of hardware/software partitioning, communications costs, and opportunities for concurrent computation between software and hardware. Analysis of these factors is beyond the scope of current automatic parallelizing compilers. In this paper, a case study is presented in which a simulation of metropolitan road traffic networks is mapped onto a reconfigurable supercomputer, the Cray XD1. Five different methods are presented for mapping the application onto the combined hardware/software system. An approach for approximating the performance of each method is derived through analytic equations. Our results, both analytically and empirically, show that key predictors of performance (which are often not considered in reported speedup of kernel operations) are not necessarily maximum parallelism, but must account for the fraction of the problem that runs on the reconfigurable logic and the amount data flow between software and hardware. Justin L. Tripp, Maya B. Gokhale, Anders A. Hansson |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | On the Acceleration of Shortest Path Calculations in Transportation NetworksabstractShortest path algorithms are key elements of many graph problems. They are used in such applications as online direction finding and navigation, and modeling of traffic for large scale simulations of major metropolitan areas. As shortest path algorithm are execution bottlenecks, it is beneficial to move their execution to parallel hardware such as field programmable gate arrays (FPGAs). One of the innovations of this approach is the use of a small bubble sort core to produce the extract-min function. While bubble sort is not usually considered an appropriate algorithm for any non-trivial usage, it is appropriate in this case as it can produce a single minimum out of the list in O(n) cycles, where n is the number of elements in the vertex list. The cost of this min operation does not impact the running time of the architecture, because the queue depth for fetching the next set of edges from memory is roughly equivalent to the number of cores in the system. Additionally, this work provides a collection of simulation results that model the behavior of the node queue in hardware. The results show that a hardware queue, implementing a small bubble-type minimum function, need only be on the order of 16 elements to provide both correct and optimal paths. With support for a large DRAM graph store with SRAM-based caching on a Cray XD-1 FPGA-accelerated system, the system provides a speedup of roughly 50x over the CPU-based implementation. Zachary K. Baker, Maya B. Gokhale |
FCCM | 2 |
| 2007 | Matched Filter Computation on FPGA, Cell and GPUabstractThe matched filter is an important kernel in the processing of hyperspectral data. The filter enables researchers to sift useful data from instruments that span large frequency bands and can produce Gigabytes of data in seconds. In this work, we evaluate the performance of a matched filter algorithm implementation on an FPGA-accelerated co-processor (Cray XD-1), the IBM Cell microprocessor, and the NVIDIA GeForce 7900 GTX GPU graphics card. We provide extensive discussion of the challenges and opportunities afforded by each platform. In particular, we explore the problems of partitioning the filter most efficiently between the host CPU and the co-processor. Using our results, we derive several performance metrics that provide the optimal solution for a variety of application situations. Zachary K. Baker, Maya B. Gokhale, Justin L. Tripp |
FCCM | 2 |
| 2006 | A hybrid framework for design and analysis of fault-tolerant architecturesabstractIt is anticipated that self assembled ultra-dense nanomemories will be more susceptible to manufacturing defects and transient faults than conventional CMOS-based memories, thus the need exists for fault-tolerant memory architectures. The development of such architectures will require intense analysis in terms of achievable performance measures- power dissipation, area, delay and reliability. In this paper, we propose and develop a hybrid automation framework, called HMAN, that aids the design and analysis of fault-tolerant architectures for nanomemories. Our framework can analyze memory architectures at two different levels of the design abstraction, namely the system and circuit levels. To the best of our knowledge, this is the first such attempt at analyzing memory systems at different levels of abstraction and then correlating the different performance measures. We also illustrate the application of our framework to self-assembled crossbar architectures by analyzing a hierarchical fault-tolerant crossbar-based memory architecture that we have developed. Debayan Bhaduri, Sandeep K. Shukla, Deji Coker, Valerie Taylor 0001, Paul S. Graham, Maya B. Gokhale |
DATE | 6 |
| 2006 | A Run-Time Re-configurable Parametric Architecture for Local Neighborhood Image ProcessingabstractWe propose a run-time re-configurable parametric architecture (fabric) for local neighborhood image processing. The proposed architecture is composed of polymorphous cells where each cell accesses neighborhood data from a local cell memory, and executes a neighborhood function sequentially. The architecture is flexible since different neighborhood functions can be implemented by rewriting a cell's software micro-code. High throughput is achieved because many cells execute concurrently. We show that for a satellite image feature extraction application, our architecture, implemented on Stratix II and Virtex 2 field programmable gate arrays, achieves similar performance, hardware resource utilization, and throughput as a fully pipelined systolic array architecture, yet offers imp roved flexibility to the developer. We compare and contrast these two architectures for their usability to the image processing community Reid B. Porter, Jan R. Frigo, Maya B. Gokhale, Christophe Wolinski, François Charot, Charles Wagner |
DSD | 3 |
| 2006 | A Programmable, Maximal Throughput Architecture for Neighborhood Image ProcessingabstractThe authors propose a run-time re-configurable architecture for local neighborhood image processing. Discussion of how the new architecture can offer improved flexibility to the developer. The authors show that for a satellite image feature extraction application, our architecture, implemented on Stratix II and Virtex 2 field programmable gate arrays, achieves similar performance, hardware resource utilization, and throughput as fully pipelined systolic array architecture Reid B. Porter, Jan R. Frigo, Maya B. Gokhale, Christophe Wolinski, François Charot, Charles Wagner |
FCCM | 3 |
| 2006 | The STAR-C Truth: Analyzing Reconfigurable Supercomputing ReliabilityabstractIn this abstract, the authors present an overview of a reliability analysis toolset, called the scalable tool for the analysis of reliable systems (STAR systems), with modules for determining the reliability of FPGA designs (STAR-circuits) and reconfigurable supercomputers (STAR-reconfigurable supercomputers Heather M. Quinn, Debayan Bhaduri, Christof Teuscher, Paul S. Graham, Maya B. Gokhale |
FCCM | 5 |
| 2006 | Hardware/Software Approach to Molecular Dynamics on Reconfigurable ComputersabstractWith advances in re configurable hardware, especially field-programmable gate arrays (FPGAs), it has become possible to use reconfigurable hardware to accelerate complex applications, such as those in scientific computing. There has been a resulting development of reconfigurable computers - computers which have both general purpose processors and reconfigurable hardware, as well as memory and high-performance interconnection networks. In this paper, we study the acceleration of molecular dynamics simulations using reconfigurable computers. We describe how we partition the application between software and hardware and then model the performance of several alternatives for the task mapped to hardware. We describe an implementation of one of these alternatives on a reconfigurable computer and demonstrate that for two real-world simulations, it achieves a 2 times speed-up over the software baseline. We then compare our design and results to those of prior efforts and explain the advantages of the hardware/software approach, including flexibility Ronald Scrofano, Maya B. Gokhale, Frans Trouw, Viktor Prasanna 0001 |
FCCM | 2 |
| 2006 | RAW keynote 1: the outer limits: reconfigurable computing in space and in orbitabstractSummary form only given. Programmable hardware offers unique opportunities for flexible control and processing on board spacecrafts and satellites. Space missions dictate stringent requirements on size, weight, power, versatility and performance of the on-board data acquisition and computing resources. Reconfigurable FPGA-based devices, with the programmability of software and speed/size approaching application-specific integrated circuits, make it possible to control and communicate with sensors, as well as process scientific data right on the spacecraft, sending only relevant information back home over the low bandwidth communications link. Challenges to computing in harsh space environments abound: vibration, thermal cycling, heat dissipation, and radiation all take their toll on space electronics. In spite of these barriers, notable experiments in reconfigurable computing for space applications are being undertaken. These include NASA's reconfigurable scalable computing project, intended for planetary rovers, cameras, and other sensors; the Queensland University FedSat and its successors, using FPGAs for near real-time image processing, communications, and navigation; and the Cibola Flight Experiment, a Los Alamos National Laboratory experiment in on-orbit signal processing using radiation-tolerant FPGAs. This talk discusses the perils and possibilities of reconfigurable computing at the outer limits. Maya B. Gokhale |
IPDPS | 1 |
| 2006 | Panel: Nano-computing - do we need new formal approaches?abstractWith current CMOS technologies reaching beyond 65 nanometers mark, and the highlights of computing fabrics such as molecular, DNA guided assemblies, quantum comput- ing, carbon nanotube based transistors etc. are bringing the focus onto nanotechnology.The term nano-computing now implies computing with alternative emerging nano-scale technologies as well as with large scale parallelism afforded by the shrinking silicon technologies, with the added burden of rampant defects and faults. The optimization factors and assumptions we made so far in hardware designs become no longer valid for such computing fabrics. The parallelism heretofore unavailable can now be applied, and help us realize newer ways of designing softwares and algorithms. The question therefore to ask is: do we need drastic new approaches to design hardware and software? This panel comprising of experts from formal methods, reliability, spatial computing, hardware/software design will discuss this question from individual perspectives in an attempt to come up with a cogent set of questions researchers need to answer. M. Hsiao, Sandeep K. Shukla, Maya B. Gokhale, Alvin R. Lebeck |
MEMOCODE | 3 |
| 2005 | Metropolitan Road Traffic Simulation on FPGAsabstractThis work demonstrates that road traffic simulation of entire metropolitan areas is possible with reconfigurable supercomputing that combines 64-bit microprocessors and FPGAs in a high bandwidth, low latency interconnect. Previously, traffic simulation on FPGAs was limited to very-short road segments or required a very large number of FPGAs. Our data streaming approach overcomes scaling issues associated with direct implementations and still allows for high-level parallelism by dividing the data sets between hardware and software across the reconfigurable supercomputer. Using one FPGA on the Cray XD1 supercomputer, we are able to achieve a 34.4/spl times/ speed up over the AMD microprocessor. System integration issues must be optimized to exploit this speedup in the overall simulation. Justin L. Tripp, Henning S. Mortveit, Anders A. Hansson, Maya B. Gokhale |
FCCM | 4 |
| 2005 | Trident: An FPGA Compiler Framework for Floating-Point AlgorithmsabstractTrident is a compiler for floating point algorithms written in C, producing circuits in reconfigurable logic that exploit the parallelism available in the input description. Trident automatically extracts parallelism and pipelines loop bodies using conventional compiler optimizations and scheduling techniques. Trident also provides an open framework for experimentation, analysis, and optimization of floating point algorithms on FPGAs and the flexibility to easily integrate custom floating point libraries. Justin L. Tripp, Kristopher D. Peterson, Christine Sweeney, Jeffrey D. Poznanovic, Maya B. Gokhale |
FPL | 5 |
| 2005 | Partitioning Hardware and Software for Reconfigurable Supercomputing Applications: A Case StudyabstractOften reconfigurable systems are reported to have 10× to 100× speedup over that of a software system. However, the reconfigurable hardware must usually be combined with software to form an entire system. This system integration presents a hardware/software co-design problem with many system engineering issues. Here, we present traffic acceleration on the Cray XD1 supercomputer and describe the costs involved in different hardware/software trade-offs. Justin L. Tripp, Anders A. Hansson, Maya B. Gokhale, Henning S. Mortveit |
SC | 3 |
| 2004 | A Constraints Programming Approach to Communication Scheduling on SoPC ArchitecturesabstractThis paper presents a method to obtain an optimized static schedule of CSP-like communications between a collection of concurrent hardware processes implemented on "system on a programmable chip" (SoPC). The hardware processes are the applications tailored "cells" in the processor-coupled polymorphous fabric (Wolinski et al., 2003 and Wolinski et al., 2002) implemented on the Altera Excalibur Arm SoPC platform. In our global approach, a static communications schedule is adopted to reduce hardware overhead due to inter-cells message transfer synchronization and to speed-up application execution. The scheduling problem is defined and solved using constraints programming approach. This approach makes it possible to obtain optimal communication schedules in a number of real cases. It makes also possible to easily generate pipelined schedules that improve significantly performance of the final implementation. Our method is illustrated with a fabric-based implementation of the K-means clustering algorithm. An optimal communication scheduled is achieved for this application. Christophe Wolinski, Krzysztof Kuchcinski, Maya B. Gokhale |
DSD | 3 |
| 2004 | Communications Scheduling for Concurrent Processes on Reconfigurable ComputersabstractWe describe a unified approach to scheduling point-to-point uni-directional communications among concurrent FPGA-based hardware processes. In this model, processes have separate address spaces, and share data through communication. Once a channel is written, it may not be re-written until the receiving process reads the data. Thus if the writer process is ready before the reader has read the previous message, the writer must stall. We present an algorithm to automatically generate synchronized hardware schedules for the parallel processes that communicate, so that hardware stall management is not required. The algorithm requires that the parallel processes conform to certain constraints in program control structures and communications forms. If the processes do not conform to these requirements, hardware-supported stall mechanisms are used. We quantify the impact in area and clock speed between compiler-generated synchronization of process schedules and run-time, hardware-mediated synchronization. Maya B. Gokhale, Christine Sweeney, Janette Frigo, Christophe Wolinski |
FCCM | 1 |
| 2004 | A constraints programming approach to communication scheduling on SoPC architecturesabstractThis paper presents a novel approach to scheduling communications among concurrent hardware processes mapped onto a "System on a Programmable Chip." Point-to-point, broadcast and multi-cast ommunication types are supported. The algorithm has been prototyped on the Processor-Coupled Polymorphous Fabric for the Altera Excalibur Arm architecture. The communication schedule problem has been specified using Constraints Programming. The advantages of our method are the following: Application of a general constraint solver makes it possible to express many different sorts of constraints in a uniform manner. All imposed constraints are handled by the solver concurrently which increases the chances of obtaining optimal results. The scheduler guarantees that loops periods are the same for each iteration so the smaller controllers can be generated. The method has been illustrated with a Fabric-based implementation of the K-means clustering algorithm for which an optimal communication schedule has been achieved. Christophe Wolinski, Krzysztof Kuchcinski, Maya B. Gokhale |
FPGA | 3 |
| 2004 | Monte Carlo Radiative Heat Transfer Simulation on a Reconfigurable Computer
Maya B. Gokhale, Janette Frigo, Christine Sweeney, Justin L. Tripp, Ron Minnich |
FPL | 1 |
| 2004 | Dynamic Reconfiguration for Management of Radiation-Induced Faults in FPGAsabstractSummary form only given. We describe novel methods of exploiting the partial, dynamic reconfiguration capabilities of Xilinx Virtex 1000 FPGAs to manage transient faults due to radiation in space environments. The on-orbit fault detection scheme uses a radiation-hardened reconfiguration controller to continuously monitor the configuration bit streams of 9 Virtex FPGAs and to correct errors by partial, dynamic reconfiguration of the FPGAs while they continue to execute. To study single event upset (SEU) impact on our signal processing applications, we use a novel fault injection technique to corrupt configuration bits, thereby simulating SEU faults. By using dynamic reconfiguration, we can run the corrupted designs directly on the FPGA hardware, giving many orders of magnitude speed-up over purely software techniques. The fault injection method has been validated against proton beam testing, showing 97.6% agreement. Our work highlights the benefits of dynamic reconfiguration for space-based reconfigurable computing. Maya B. Gokhale, Paul S. Graham, Darrel Eric Johnson, Nathan Rollins, Michael J. Wirthlin |
IPDPS | 1 |
| 2003 | Gamma-Ray Pulsar Detection using Reconfigurable Computing HardwareabstractThis paper presents a method to detect gamma-ray pulsars using a fast folding algorithm (Staelin, 1969) mapped onto reconfigurable hardware. In contrast, existing techniques require gigapoint complex FFTs. the algorithm has been written in Streams-C and compiled with the sc2 compiler to the target Annapolis Micro Systems (AMS) Firebird board (Xilinx Virtex E processor). To accelerate detection of new gamma-ray pulsars, the sc2 compiler generates a hardware implementation of the algorithm for finding periodicities in data sets. The data to be analyzed comes from a high-energy gamma-ray telescope onboard a spacecraft. This astrophysics application poses a "good example" of the use of a high level reconfigurable computing tool such as sc2 to accelerate an algorithm because it uses real satellite data, the algorithm can be parallelized, and was originally validated using a high level scientific language, IDL. By recasting the algorithm into Streams-C, the scientific software developer can create a hardware implementation on a reconfigurable computing platform. We describe the fast folding algorithm, the Streams-C implementation, and discuss techniques to optimize performance within the Streams-C framework. The compiler-generated hardware delivers approximately 3X to 6X speed up over a comparable 800MHz general-purpose processor doing the software-only algorithm. Janette Frigo, David Palmer 0006, Maya B. Gokhale, Marc Popkin-Paine |
FCCM | 3 |
| 2003 | Fabric-Based Systems: Model, Tools, ApplicationsabstractA Fabric-Based System (FBS) is a parameterized cellular architecture in which an array of computing cells communicates with an embedded processor through a global memory. This architecture is customizable to different classes of applications by functional unit, interconnect, and memory parameters, and can be instantiated efficiently on platform FPGAs (field programmable gate array). In previous work (Bellows et al., 2002), we have demonstrated the advantage of reconfigurable fabrics for image and signal processing applications. Recently, we have built a Fabric Generator (FG), a Java-based toolset that greatly accelerates construction of the fabrics. A module-generation library is used to define, instantiate, and interconnect cells' datapaths. FG also generates customized sequencers for individual cells or collections of cells. We describe the Fabric-Based System model, the FG toolset, and concrete realizations of fabric architectures generated by FG on the Altera Excalibur ARM that can deliver 4.5 GigaMACs/s (8/16 bit data, multiply-accumulate). Christophe Wolinski, Maya B. Gokhale, Kevin McCabe |
FCCM | 2 |
| 2003 | Polymorphous fabric-based systems: Model, tools, applications
Christophe Wolinski, Maya B. Gokhale, Kevin McCabe |
J. Syst. Archit. | 2 |
| 2003 | Experience with a Hybrid Processor: K-Means Clustering
Maya B. Gokhale, Janette Frigo, Kevin McCabe, James Theiler, Christophe Wolinski, Dominique Lavenier |
J. Supercomput. | 1 |
| 2002 | Granidt: Towards Gigabit Rate Network Intrusion Detection Technology
Maya B. Gokhale, Dave Dubois, Andy Dubois, Mike Boorman, Stephen W. Poole, Vic Hogsett |
FPL | 1 |
| 2001 | Mutable Functional Units: Initial Results
Yan Solihin, Kirk W. Cameron, Dominique Lavenier, Maya B. Gokhale |
FCCM | 5 |
| 2001 | Evaluation of the streams-C C-to-FPGA compiler: an applications perspectiveabstractThe Streams-C compiler ([5]) synthesizes hardware circuits for reconfigurable FPGA-based computers from parallel C programs. The Streams-C language consists of a small number of libraries and intrinsic functions added to a synthesizable subset of C, and supports a communicating process programming model. The processes may be either software or hardware processes, and the compiler manages communication among the processes transparently to the programmer. For the hardware processes, the compiler generates Register-Transfer-Level (RTL) VHDL, targeting multiple FPGAs with dedicated memories. For the software processes, a multi-threaded software program is generated. Janette Frigo, Maya B. Gokhale, Dominique Lavenier |
FPGA | 2 |
| 2001 | Mutable Functional Units and Their Applications on MicroprocessorsabstractFunctional units are the heart of microprocessors as they execute binary instructions of a program. Current microprocessors typically have several types of functional units. In this paper, we propose a new functional unit that combines a floating-point adder and an integer arithmetic and logic unit into a single unit. This functional unit reconfigures itself at run-time to serve different instructions from the program instruction stream. We call such units mutable functional units or MFUs. MFUs can be used in microprocessors to improve functional unit utilization, reduce power consumption, and to improve performance without adding extra functional units. MFUs only require, minor modifications to the existing floating-point adder design. We show that overheads of reconfiguration are small, typically 0 to 1 clock cycle, and at most 2 clock cycles. We demonstrate how integration with a typical current microprocessor can be achieved. This integration allows speedups of non-numerical applications by 8% to 14% while keeping the number of functional units constant. We also show that various enhancements to the base architecture that increase the instruction fetch rate affect the speedups positively. Yan Solihin, Kirk W. Cameron, Dominique Lavenier, Maya B. Gokhale |
ICCD | 5 |
| 2000 | Stream-Oriented FPGA Computing in the Streams-C High Level LanguageabstractStream oriented processing is an important methodology used in FPGA-based parallel processing. Characteristics of stream-oriented computing include high-data-rate flow of one or more data sources; fixed size, small stream payload (one byte to one word); compute-intensive operations, usually low precision fixed point, on the data stream; access to small local memories holding coefficients and other constants; and occasional synchronization between computational phases. We describe language constructs, compiler technology, and hardware/software libraries embodying the Streams-C system which has been developed to support stream-oriented computation on FPGA-based parallel computers. The language is implemented as a small set of library functions callable from a C language program. The Streams-C compiler synthesizes hardware circuits for multiple FPGAs as well as a multi-threaded software program for the control processor. Our system includes a functional simulation environment based on POSIX threads, allowing the programmer to simulate the collection of parallel processes and their communication at the functional level. Finally we present an application written both in Streams-C and hand-coded in VHDL. Compared to the hand-crafted design, the Streams-C-generated circuit takes 3x the area and runs at 1/2 the clock rate. In terms of time to market, the hand-done design took a month to develop by an experienced hardware developer. The Streams-C design rook a couple of days, for a productivity increase of 10x. Maya B. Gokhale, Janice M. Stone, Jeffrey M. Arnold, Mirek Kalinowski |
FCCM | 1 |
| 1999 | Automatic Allocation of Arrays to Memories in FPGA Processors with Multiple Memory BanksabstractFPGA-based processors, like many conventional DSP systems, often associate small high performance memories with each processing chip. These memories may be on-board embedded SRAMs or discrete parts. In the process of mapping a computation onto an FPGA processor, it is necessary to map the applications' data to memories. In this work, we present an algorithm that has been implemented in our NAPA C compiler to assign data automatically to memories to produce minimum overall execution time of the loops in the program. With the addition of this algorithm to our compiler, the programmer need not explicitly annotate array declarations with location information. Rather, the compiler analyzes the usage patterns of variables and selects the optimal location for each variable. The algorithm uses a search technique known as implicit enumeration to reduce the otherwise exponential search space. In practice, the use of this memory allocation compiler phase in our SUIF-based NAPA C compiler has negligible effect on compiler run time. Maya B. Gokhale, Janice M. Stone |
FCCM | 1 |
| 1998 | NAPA C: Compiling for a Hybrid RISC/FPGA ArchitectureabstractHybrid architectures combining conventional processors with configurable logic resources enable efficient coordination of control with datapath computation. With integration of the two components on a single device, loop control and data-dependent branching can be handled by the conventional processor. While regular datapath computation occurs on the configurable hardware. This paper describes a novel pragma-based approach to programming such hybrid devices. The NAPA C language provides pragma directives so that the programmer (or an automatic partitioner) can specify where data is to reside and where computation is to occur with statement-level granularity. The NAPA C compiler, targeting National Semiconductor's NAPA1000 chip, performs semantic analysis of the pragma-annotated program and co-synthesizes a conventional program executable combined with a configuration bit stream for the adaptive logic. Compiler optimizations include synthesis of hardware pipelines from pipelineable loops. Maya B. Gokhale, Janice M. Stone |
FCCM | 1 |
| 1998 | The NAPA Adaptive Processing ArchitectureabstractThe National Adaptive Processing Architecture (NAPA) is a major effort to integrate the resources needed to develop teraops class computing systems based on the principles of adaptive computing. The primary goals for this effort include: (1) the development of an example NAPA component which achieves an order of magnitude cost/performance improvement compared to traditional FPGA based systems, (2) the creation of a rich but effective application development environment for NAPA systems based on the ideas of compile time functional partitioning and (3) significantly improve the base infrastructure for effective research in reconfigurable computing. This paper emphasizes the technical aspects of the architecture to achieve the first goal while illustrating key architectural concepts motivated by the second and third goals. Charlé R. Rupp, Mark Landguth, Tim Garverick, Edson Gomersall, Harry Holt, Jeffrey M. Arnold, Maya B. Gokhale |
FCCM | 7 |
| 1997 | High level compilation for fine grained FPGAsabstractThe authors present an integrated tool set to generate highly optimized hardware computation blocks from a C language subset. By starting with a C language description of the algorithm, they address the problem of making FPGA processors accessible to programmers as opposed to hardware designers. Their work is specifically targeted to fine grained FPGAs such as the National Semiconductor CLAy/sup TM/ FPGA family. Such FPGAs exhibit extremely high performance on regular data path circuits, which are more prevalent in computationally oriented hardware applications. Dense packing of data path functional elements makes it possible to fit the computation on one or a small number of chips, and the use of local routing resources makes it possible to clock the chip at a high rate. By developing a lower level tool suite that exploits the regular, geometric nature of fine grained FPGAs, and mapping the compiler output to this tool suite, they greatly improve performance over traditional high level synthesis to fine grained FPGAs. Maya B. Gokhale, D. Gomersall |
FCCM | 1 |
| 1995 | Data-parallel C on a reconfigurable logic array
Maya B. Gokhale, Brian Schott |
J. Supercomput. | 1 |
| 1993 | SIMD Optimizations in a Data Parallel CabstractSIMD programs can devote substantial time to manipulating the underlying hardwares context registers status bits that determine whether processors in the SIMD array execute or skip the current instruction. This paper describes two optimizations, implemented in a compiler for a data parallel C, that reduce the overhead of manipulating and accessing context registers. The first optimization uses observations about a program's nesting structure to eliminate context register save/restore operations performed by guarded parallel control constructs. The second uses two-version code to eliminate context register acceaaes performed by individual instructions. Maya B. Gokhale, Phil Pfeiffer |
ICPP (2) | 1 |
| 1992 | An introduction to compilation issues for parallel machines
Maya B. Gokhale, William Carlson |
J. Supercomput. | 1 |
| 1992 | Parallel Evaluation of Attribute GrammarsabstractExamines the generation of parallel evaluators for attribute grammars, targeted to shared-memory MIMD computers. Evaluation-time overhead due to process scheduling and synchronization is reduced by detecting coarse-grain parallelism (as opposed to the naive one-process-per-node approach). As a means to more clearly expose inherent parallelism, it is shown how to automatically transform productions of the form X to Y X into list-productions of the form X to Y/sup +/. This transformation allows for many simplifications to be applied to the semantic rules, which can expose a significant degree of inherent parallelism, and thus further increase the evaluator's performance. Effectively, this constitutes an extension of the concept of attribute grammars to the level of abstract syntax.> Alexander C. Klaiber, Maya B. Gokhale |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1990 | The Logic Description GeneratorabstractThe authors describe the Logic Description Generator (LDG), a design tool specifically geared to aid in the implementation of systolic algorithms on reconfigurable logic arrays. It is used to specify designs for Splash, a linear array of Xilinx chips. LDG supports the notion of a logical systolic cell, which may be repetitively layed out across a chip, and whose instances may be interconnected as a linear array. LDG also contains a reshape operator, which allows the hardware designer to place each component of the linear array at a specific location on the chip. Another feature of LDG is the ability to write general parameterized library routines. LDG allows the designer to make a simple change to a single cell and then have that change affect the configuration and layout of every other cell on the chip in a time frame comparable to the software edit-compile-test cycle. The turnaround time from textual chip specification to testing the new configuration on the Splash hardware is on the order of a half-hour, of which the LDG process contributes about five minutes. LDG is implemented in Common Lisp and runs in the Sun workstation environment. The LDG processor generates Xilinx Netlist format (XNF), which is a textual description of each logic element on the chip.> Maya B. Gokhale, Andrew Kopser, Sara P. Lucas, Ron Minnich |
ASAP | 1 |
| 1990 | SPLASH: A Reconfigurable Linear Logic Array
Maya B. Gokhale, William Holmes, Andrew Kopser, Dick Kunze, Daniel P. Lopresti, Sara P. Lucas, Ron Minnich, Peter Olsen |
ICPP (1) | 1 |
| 1989 | Parallel Evaluation of Attribute Grammars
Alexander C. Klaiber, Maya B. Gokhale |
ICPP (3) | 2 |
| 1988 | The symbolic hyperplane transformation for recursively defined arraysabstractA restructuring transformation that can be used to parallelize recurrence relations is described. The transformation is based on the hyperplane (or wavefront) method, but extends the applicability of the method to irregularly structured recurrences. It is assumed that the time of computation of an array element is a linear combination of its indices, and a variation of the simplex method is used to seek a succession of hyperplanes along which array elements can be computed concurrently. Portions of the work reported here have been implemented in the PS automatic program generation system.> Maya B. Gokhale, Todd C. Torgersen |
SC | 1 |
| 1988 | Parallel Scheduling of Recursively Defined Arrays
Thomas J. Myers, Maya B. Gokhale |
J. Symb. Comput. | 2 |
| 1987 | Exploiting Loop Level Parallelism in Nonprocedural Dataflow Programs
Maya B. Gokhale |
ICPP | 1 |
| 1986 | Macro vs. Micro Dataflow: A Programming Example
Maya B. Gokhale |
ICPP | 1 |