VLDB 2026 Research / reviewers in the wild / expert
Ali Akoglu
dblp:63/2031
· DBLP profile ↗
51ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0001-7982-8991ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 42 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A portable framework with generalized runtime features for task graph execution and concurrent multi-application deployment on heterogeneous systems
Serhan Gener, Md Sahil Hassan, Hasan Umut Suluhan, Liangliang Chang, Chaitali Chakrabarti, Tsung-Wei Huang, Ümit Y. Ogras, Ali Akoglu |
Future Gener. Comput. Syst. | 8 |
| 2026 | CADAS: Communication-Aware Dynamic Scheduler on CGRAs for Large-Volume and Real-Time ProcessingabstractModern data-intensive applications demand accelerators that can adapt to dynamic and high-throughput workloads. Coarse-Grained Reconfigurable Arrays (CGRAs) have emerged as promising candidates for such workloads due to their spatial architecture and run-time reconfigurability. However, ad-hoc hardware configurations and traditional static compilation techniques struggle to cope with the run-time irregularity and control-flow dynamism. This article first presents a systematic design space exploration (DSE) to identify the optimized hardware configurations tailored to application-specific constraints, such as area budget, throughput requirement, and throughput efficiency. Then, it proposes a communication-aware dynamic scheduling approach built on a hardware/software co-design that combines preloading and scoreboard mechanisms to minimize reconfiguration overhead while maximizing interconnect bandwidth utilization. Evaluated on the optimized configurations and the respective spectrum sensing benchmarks, the proposed scheduling method achieves up to 1.6× performance improvement over a baseline and 1.3× over an adapted state-of-the-art (SOTA) dynamic scheduling strategy. Hasan Umut Suluhan, Chaitali Chakrabarti, Ali Akoglu, Ümit Y. Ogras |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | Tutorial: CEDR: A Holistic Software and Hardware Design Environment for Hardware Agnostic Application Development and Deployment on FPGA-Integrated Heterogeneous SystemsabstractThis hands-on tutorial introduces CEDR, an open-source compilation and runtime framework designed to simplify application development and deployment on heterogeneous computing platforms composed of general purpose processors and hardware accelerators. CEDR unifies programming and execution on such platforms, enabling streamlined development workflows. Building on prior lectures and tutorials at ESWEEK 2023 and ISFPGA 2023/2025, this session aims to lower the barrier to research and prototyping in heterogeneous systems. Tailored for application developers and system architects, the tutorial in collaboration with the AMD University Program focuses on key capabilities of CEDR through hands-on programming exercises, including application development, functional verification, performance analysis, and design space exploration using emulation on FPGA-based SoCs. Serhan Gener, Md Sahil Hassan, Ali Akoglu |
CODES+ISSS | 3 |
| 2025 | K-PACT: Kernel Planning for Adaptive Context Switching - A Framework for Clustering, Placement, and Prefetching in Spectrum SensingabstractEfficient wideband spectrum sensing requires rapid evaluation and re-evaluation of signal presence and type across multiple subchannels. These tasks involve multiple hypothesis testing, where each hypothesis is implemented as decision tree workflow with compute-intensive kernels, including FFT, matrix operations, and signal-specific analyses. Given the dynamic nature of the spectrum environment, the ability to quickly switch between hypotheses is essential for maintaining low-latency, high-throughput operation. This work assumes a coarse-grained reconfigurable architecture consisting of an array of processing elements (PEs), each equipped with a local instruction memory (IMEM) capable of storing and executing kernels used in spectrum sensing applications. We propose a planner tool that efficiently maps hypothesis workflows onto this architecture to enable fast runtime context switching with minimal overhead. The planner performs two key tasks: clustering temporally non-overlapping kernels to share IMEM resources within a PE sub-array, and placing these clusters onto hardware to ensure efficient scheduling and data movement. By preloading kernels that are not simultaneously active into the same IMEM, our tool enables low-latency reconfiguration without runtime conflicts. It models the planning process as a multi-objective optimization, balancing trade-offs among context switch overhead, scheduling latency, and dataflow efficiency. We evaluate the proposed tool in simulated spectrum sensing scenario with 48 concurrent subchannels. Results show that our approach reduces off-chip binary fetches by 207.81×, lowers average switching time by 98.24×, and improves per-subband execution time by 132.92× over baseline without preloading. These improvements demonstrate that intelligent planning is critical for adapting to fast-changing spectrum environments in next-generation radio frequency systems. Hasan Umut Suluhan, Serhan Gener, Chaitali Chakrabarti, Ümit Y. Ogras, Ali Akoglu |
ICCAD | 6 |
| 2025 | Coarse-Grained Task Parallelization by Dynamic Profiling for Heterogeneous SoC-Based Embedded SystemabstractIn this study, we introduce a methodology for automatically transforming user applications written in C/C++ to a parallel representation consisting of coarse-grained tasks based on dynamic profiling. Such a parallel representation is suitable for mapping applications onto heterogeneous SoCs. We present our approach for instrumenting the user application binary during the compilation process with parallel primitives that enable the runtime system to schedule and execute independent computation-intensive coarse-grained tasks concurrently. We use the proposed compilation and code transformation methodology to retarget each application for execution on a heterogeneous SoC composed of processor cores and accelerators. We demonstrate the capabilities of our integrated compile time and runtime flow through task-level parallelization and functionally correct execution of real-world applications in the communication systems and radar processing domains. We demonstrate the functionality of our integrated system by executing six distinct applications with different degrees of parallelism on four different platforms: an eight-core general-purpose processor, a heterogeneous SoC simulator, and two heterogeneous SoCs utilizing the Xilinx Zynq UltraScale+ FPGA and the Nvidia Jetson AGX board. Our integrated approach offers a path forward for application developers to take full advantage of the target SoC without requiring users to become hardware or parallel programming experts. Liangliang Chang, Serhan Gener, Joshua Mack, Hasan Umut Suluhan, Ali Akoglu, Chaitali Chakrabarti |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2025 | RIMMS: Runtime Integrated Memory Management System for Heterogeneous ComputingabstractEfficient memory management in heterogeneous systems is increasingly challenging due to diverse compute architectures (e.g., CPU, GPU, and FPGA) and dynamic task mappings not known at compile time. Existing approaches often require programmers to manage data placement and transfers explicitly, or assume static mappings that limit portability and scalability. This article introduces RIMMS (Runtime Integrated Memory Management System), a lightweight, runtime-managed, hardware-agnostic memory abstraction layer that decouples application development from low-level memory operations. RIMMS transparently tracks data locations, manages consistency, and supports efficient memory allocation across heterogeneous compute elements without requiring platform-specific tuning or code modifications. We integrate RIMMS into a baseline runtime and evaluate with complete radar signal processing applications across CPU+GPU and CPU+FPGA platforms. RIMMS delivers up to 2.43× speedup on GPU-based and 1.82× on FPGA-based systems over the baseline. Compared to IRIS, a recent heterogeneous runtime system, RIMMS achieves up to 3.08X speedup and matches the performance of native CUDA implementations while significantly reducing programming complexity. Despite operating at a higher abstraction level, RIMMS incurs only 1–2 cycles of overhead per memory management call, making it a low-cost solution. These results demonstrate RIMMS’s ability to deliver high performance and enhanced programmer productivity in dynamic, real-world heterogeneous environments. Serhan Gener, Aditya Ukarande, Shilpa Mysore Srinivasa Murthy, Md Sahil Hassan, Joshua Mack, Chaitali Chakrabarti, Ümit Y. Ogras, Ali Akoglu |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2025 | Tutorial: A Novel Runtime Environment for Accelerator-Rich Heterogeneous ArchitecturesabstractAs the landscape of computing advances, system designers are increasingly exploring methodologies that leverage higher levels of heterogeneity to enhance performance within constrained size, weight, power, and cost parameters. CEDR (Compiler-integrated Extensible DSSoC Runtime) stands as an ecosystem facilitating productive and efficient application development and deployment across heterogeneous computing systems. It fosters the co-design of applications, scheduling heuristics, and accelerators within a unified framework. Our goal is to present CEDR as a promising environment for lifting the barriers to research on heterogeneous systems and addressing the broader challenges within domain-specific architectures. We introduce CEDR and discuss the evolutionary design decisions underlying its programming model. Subsequently, we explore its utility for a broad range of users through design sweeps on off-the-shelf heterogeneous platforms across scheduling heuristics, hardware compositions, and workload scenarios. Joshua Mack, Anish Krishnakumar, Ümit Y. Ogras, Ali Akoglu |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2023 | PyTorch and CEDR: Enabling Deployment of Machine Learning Models on Heterogeneous Computing SystemsabstractThe PyTorch programming interface enables efficient deployment of machine learning models, leveraging the parallelism offered by GPU architectures. In this study, we present the integration of the PyTorch framework with a compiler and runtime ecosystem . Our aim is to demonstrate the ability to deploy PyTorch-based models on FPGA-based SoC platforms, without requiring users to possess prior FPGA-based design experience. The proposed PyTorch model transformation approach expands the range of hardware architectures that PyTorch developers can target, enabling them to take advantage of the energy-efficient execution provided by heterogeneous computing systems. Our experiments involve compiling and executing real-life applications on heterogeneous SoC configurations emulated on the Xilinx Zynq Ultrascale+ ZCU102 system. We showcase our ability to deploy three distinct PyTorch applications, encompassing object detection, visual geometry group (VGG), and speech classification, using the integrated compiler and runtime system without loss of model accuracy. Furthermore, we extend our analysis by evaluating dynamically arriving workload scenarios, consisting of a mix of PyTorch models and non-PyTorch-based applications. Through these experiments, we vary the hardware composition and scheduling heuristics. Our findings indicate that when PyTorch-based applications coexist with unrelated applications, our integrated scheduler fairly dispatches tasks to the FPGA platform’s accelerator and CPU cores, without compromising the target throughput for each application. Hasan Umut Suluhan, Serhan Gener, Alexander Fusco, H. Fatih Ugurdag, Ali Akoglu |
AICCSA | 5 |
| 2023 | A Novel Implementation Methodology for Error Correction Codes on a Neuromorphic ArchitectureabstractThe Internet of Things infrastructure connects a massive number of edge devices with an increasing demand for intelligent sensing and inferencing capability. Such data-sensitive functions necessitate energy-efficient and programmable implementations of error correction codes (ECCs) and decoders. The algorithmic flow of ECCs with concurrent accumulation and comparison types of operations are innately exploitable by neuromorphic architectures for energy-efficient execution—an area that is relatively unexplored outside of machine learning applications. For the first time, we propose a methodology to map the hard-decision class of decoder algorithms on a neuromorphic architecture. We present the implementation of the Gallager B (GaB) decoding algorithm on a TrueNorth-inspired architecture that is emulated on the Xilinx Zynq ZCU102 MPSoC. Over this reference implementation, we propose architectural modifications at the neuron block level that result in a reduction of energy consumption by 31% with a negligible increase in resource usage while achieving the same error correction performance. Md Sahil Hassan, Parker Dattilo, Ali Akoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | CEDR: A Compiler-integrated, Extensible DSSoC RuntimeabstractIn this work, we present a C ompiler-integrated, E xtensible D omain Specific System on Chip R untime (CEDR) ecosystem to facilitate research toward addressing the challenges of architecture, system software, and application development with distinct plug-and-play integration points in a unified compile time and runtime workflow. We demonstrate the utility of CEDR on the Xilinx Zynq MPSoC-ZCU102 for evaluating performance of pre-silicon hardware in the trade space of SoC configuration, scheduling policy and workload complexity based on dynamically arriving workload scenarios composed of real-life signal processing applications scaling to thousands of application instances with Fast Fourier Transform and matrix multiply accelerators. We provide insights into the tradeoffs present in this design space through a number of distinct case studies. CEDR is portable and has been deployed and validated on Odroid-XU3, X86, and Nvidia Jetson Xavier-based SoC platforms. Taken together, CEDR is a capable environment for enabling research in exploring the boundaries of productive application development, resource management heuristic development, and hardware configuration analysis for heterogeneous architectures. Joshua Mack, Md Sahil Hassan, Nirmal Kumbhare, Miguel Castro-Gonzalez, Ali Akoglu |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2022 | GPGPU-based High Throughput Image Pre-processing Towards Large-Scale Optical Character RecognitionabstractStudies have shown that pre-processing digital images through scaling, rotation and blurring type of operations allow optical character recognition (OCR) to focus on the key features in the image and result in improving recognition accuracy. We leverage the open-source Tesseract OCR and show that its accuracy can be improved through a pre-processing flow that includes thresholding, rotation, rescaling, erosion, dilation, and noise removal steps based on a dataset that is formed of 560 phone screen images. However, the serial CPU-based implementation of this flow introduces a latency of 48.32 ms per image on average. Even though time scale is low in the context of a single image, this latency poses as a barrier when processing millions of images with OCR. To address this, we parallelize the entire pre-processing flow on the Nvidia P100 GPU, implement a streaming based execution, and reduce the latency to 0.846 ms. This streaming-enabled implementation enables setting up a GPU based OCR engine to process large scale workloads. Serhan Gener, Parker Dattilo, Dhruv Gajaria, Alexander Fusco, Ali Akoglu |
AICCSA | 5 |
| 2022 | Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous ProcessorabstractRF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design. Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue |
ISCAS | 3 |
| 2022 | A Hardware-based HEFT Scheduler Implementation for Dynamic Workloads on Heterogeneous SoCsabstractNon-uniform performance and power consumption across the processing elements (PEs) of heterogeneous SoCs increase the computation complexity of the task scheduling problem compared to homogeneous architectures. Latency of a software-based scheduler with the increased heterogeneity level in terms of number and types of PEs creates the necessity of deploying a scheduler as an overlay processor in hardware to be able to make scheduling decisions rapidly and enable deployment of real-life applications on heterogeneous SoCs. In this study we present the design trade-offs involved for implementing and deploying the runtime variant of the heterogeneous earliest finish time algorithm (HEFTRT) on the FPGA. We conduct performance evaluations on an SoC configuration emulated over the Xilinx Zynq ZCU102 platform. In a runtime environment we demonstrate hardware-based HEFTRT’s ability to make scheduling decisions with 9.144 ns latency on average, process 26.7% more tasks per second compared to its software counterpart, and reduce the scheduling latency by up to a factor of 183 based on workloads composed of a mixture of dynamically ×arriving real-life signal processing applications. Alexander Fusco, Md Sahil Hassan, Joshua Mack, Ali Akoglu |
VLSI-SoC | 4 |
| 2022 | JITA4DS: Disaggregated Execution of Data Science Pipelines Between the Edge and the Data CentreabstractThis paper targets the execution of data science (DS) pipelines supported by data processing, transmission and sharing across several resources executing greedy processes. Current data science pipelines environments provide various infrastructure services with computing resources such as general-purpose processors (GPP), Graphics Processing Units (GPUs), Field Programmable Gate Arrays (FPGAs) and Tensor Processing Unit (TPU) coupled with platform and software services to design, run and maintain DS pipelines. These one-fits-all solutions impose the complete externalization of data pipeline tasks. However, some tasks can be executed in the edge, and the backend can provide just in time resources to ensure ad-hoc and elastic execution environments.This paper introduces an innovative composable “Just in Time Architecture” for configuring DCs for Data Science Pipelines (JITA-4DS) and associated resource management techniques. JITA-4DS is a cross-layer management system that is aware of both the application characteristics and the underlying infrastructures to break the barriers between applications, middleware/operating system, and hardware layers. Vertical integration of these layers is needed for building a customizable Virtual Data Center (VDC) to meet the dynamically changing data science pipelines’ requirements such as performance, availability, and energy consumption. Accordingly, the paper shows an experimental simulation devoted to run data science workloads and determine the best strategies for scheduling the allocation of resources implemented by JITA-4DS. Genoveva Vargas-Solar, Md Sahil Hassan, Ali Akoglu |
J. Web Eng. | 3 |
| 2022 | Performant, Multi-Objective Scheduling of Highly Interleaved Task Graphs on Heterogeneous System on Chip DevicesabstractPerformance-, power-, and energy-aware scheduling techniques play an essential role in optimally utilizing processing elements (PEs) of heterogeneous systems. List schedulers, a class of low-complexity static schedulers, have commonly been used in static execution scenarios. However, list schedulers are not suitable for runtime decision making, particularly when multiple concurrent applications are interleaved dynamically. For such cases, the static task execution times and expectation of idle PEs assumed by list schedulers lead to inefficient system utilization and poor performance. To address this problem, we present techniques for optimizing execution of list scheduling algorithms in dynamic runtime scenarios via a family of algorithms inspired by the well-known heterogeneous earliest finish time (HEFT) list scheduler. Through dynamically arriving, realistic workload scenarios that are simulated in an open-source discrete event heterogeneous SoC simulator, we exhaustively evaluate each of the proposed algorithms across two SoCs modeled after the Xilinx Zynq Ultrascale+ ZCU102 and O-Droid XU3 development boards. Altogether, depending on the chosen variant in this family of algorithms, we are able to achieve an up to 39% execution time improvement, up to 7.24x algorithmic speedup, or up to 30% energy consumption improvement compared to the baseline HEFT implementation. Joshua Mack, Samet E. Arda, Ümit Y. Ogras, Ali Akoglu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | RANC: Reconfigurable Architecture for Neuromorphic ComputingabstractNeuromorphic architectures have been introduced as platforms for energy-efficient spiking neural network execution. The massive parallelism offered by these architectures has also triggered interest from nonmachine learning application domains. In order to lift the barriers to entry for hardware designers and application developers, we present RANC: a reconfigurable architecture for neuromorphic computing, an opensource highly flexible ecosystem that enables rapid experimentation with neuromorphic architectures in both software via C++ simulation and hardware via FPGA emulation. We present the utility of the RANC ecosystem by showing its ability to recreate behavior of IBM’s TrueNorth and validate with a direct comparison to IBM’s Compass simulation environment and published literature. RANC allows optimizing architectures based on application insights as well as prototyping future neuromorphic architectures that can support new classes of applications entirely. We demonstrate the highly parameterized and configurable nature of RANC by studying the impact of architectural changes on improving application mapping efficiency with quantitative analysis based on Alveo U250 FPGA. We present post routing resource usage and throughput analysis across implementations of synthetic aperture radar classification and vector matrix multiplication applications, and demonstrate a neuromorphic architecture that scales to emulating 259K distinct neurons and 73.3M distinct synapses. Joshua Mack, Ruben Purdy, Kris Rockowitz, Michael Inouye, Edward Richter, Spencer Valancius, Nirmal Kumbhare, Md Sahil Hassan, Kaitlin Lindsay Fair, John Mixter, Ali Akoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 11 |
| 2020 | FPGA Based High-Throughput Real-Time Feature Extraction for Modulation ClassificationabstractThe spectral correlation density (SCD) function is a feature extraction method used in signal classification systems. Due to its computational complexity, SCD has not been a desirable method for systems under power and real-time constraints. In this study, we present results for a hardware implementation of key kernels of the SCD function on a Field Programmable Gate Array (FPGA). By analyzing profiling results for a state of the art GPU implementation, we developed a preliminary architecture that is able to accelerate the most computationally demanding aspects of the SCD algorithm. We find that this FPGA architecture is able to achieve a 2.03X speedup relative to state of the art GPU-based SCD implementations by coupling SCD's large-scale data-parallel nature with an architecture well suited for fine-grained control flow and data access patterns. Joshua Mack, Ali Akoglu |
FCCM | 2 |
| 2020 | Dynamic power management for value-oriented schedulers in power-constrained HPC system
Nirmal Kumbhare, Ali Akoglu, Aniruddha Marathe, Salim Hariri, Ghaleb Abdulla |
Parallel Comput. | 2 |
| 2020 | DS3: A System-Level Domain-Specific System-on-Chip Simulation FrameworkabstractHeterogeneous systems-on-chip (SoCs) are highly favorable computing platforms due to their superior performance and energy efficiency potential compared to homogeneous architectures. They can be further tailored to a specific domain of applications by incorporating processing elements (PEs) that accelerate frequently used kernels in these applications. However, this potential is contingent upon optimizing the SoC for the target domain and utilizing its resources effectively at runtime. To this end, system-level design - including scheduling, power-thermal management algorithms and design space exploration studies - plays a crucial role. This article presents a system-level domain-specific SoC simulation (DS3) framework to address this need. DS3 enables both design space exploration and dynamic resource management for power-performance optimization of domain applications. We showcase DS3 using six real-world applications from wireless communications and radar processing domain. DS3, as well as the reference applications, is shared as open-source software to stimulate research in this area. Samet E. Arda, Anish Krishnakumar, A. Alper Goksoy, Nirmal Kumbhare, Joshua Mack, Anderson Luiz Sartor, Ali Akoglu, Radu Marculescu, Ümit Y. Ogras |
IEEE Trans. Computers | 7 |
| 2020 | A Value-Oriented Job Scheduling Approach for Power-Constrained and Oversubscribed HPC SystemsabstractIn this article, we investigate limitations in the traditional value-based algorithms for a power-constrained HPC system and evaluate their impact on HPC productivity. We expose the trade-off between allocating system-wide power budget uniformly and greedily under different system-wide power constraints in an oversubscribed system. We experimentally demonstrate that, under the tightest power constraint, the mean productivity of the greedy allocation is 38 percent higher than the uniform allocation whereas, under the intermediate power constraint, the uniform allocation has a mean productivity of 6 percent higher than the greedy allocation. We then propose a new algorithm that adapts its behavior to deliver the combined benefits of the two allocation strategies. We design a methodology with online retraining capability to create application-specific power-execution time models for a class of HPC applications. These models are used in predicting the execution time of an application on the available resources at the time of making scheduling decisions in the power-aware algorithms. We evaluate the proposed algorithm using emulation and simulation environments, and show that our adaptive strategy results in improving HPC resource utilization while delivering a mean productivity that is almost the same as the best performing algorithm across various system-wide power constraints. Nirmal Kumbhare, Aniruddha Marathe, Ali Akoglu, Howard Jay Siegel, Ghaleb Abdulla, Salim Hariri |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Accelerated Shadow Detection and Removal MethodabstractShadows can have a negative effect on the ability of computer vision techniques for object detection, tracking, and recognition. Therefore, ability to remove shadows and byproducts of illumination is an important problem to enable effective object recognition actions. As applications move into levels of higher information extraction and higher required processing speeds, efficient and sophisticated shadow detection and removal becomes even more necessary. In this study we propose a shadow removal method, parallelize using a Tesla P100 GPU, and achieve a speedup of 21.67× on an 18 megapixel (MP) resolution image compared to the same method implemented in Matlab. Edward Richter, Ryan Raettig, Joshua Mack, Spencer Valancius, Burak Unal, Ali Akoglu |
AICCSA | 6 |
| 2019 | Bit-Wise and Multi-GPU Implementations of the DNA Recombination AlgorithmabstractThe V(D)J recombination is the primary mechanism for generating a diverse repertoire of T-cell receptors (TCRs) essential to the adaptive immune system for recognizing a wide variety of diseases. However, modeling the TCR repertoire is computationally challenging as the total number of TCRs to be generated and processed can exceed 1018sequences. We propose a bit-wise implementation of the V(D) J recombination algorithm, which reduces the memory footprint and execution time by factors of 4 and 2, respectively, compared to the state-of-the-art GPU implementation. We also present a multi-GPU implementation, experimentally identify suitable workload partitioning strategies for both single-and multi-GPU implementations, and, finally, expose the relationship between the workload size and limited scalability offered by the algorithm on a cluster with up to eight GPUs. We show that the bit-wise implementation reduces the execution time from 40.5 hours to 19 hours on a single GPU and 4.4 hours on an eight-GPU configuration. Elnaz Tavakoli Yazdi, Ankur Limaye, Ali Akoglu, Tosiron Adegbija, Adam Buntzman |
HiPC | 3 |
| 2019 | Adaptive Power Reallocation for Value-Oriented Schedulers in Power-Constrained HPCabstractIn the exascale era, HPC systems are expected to operate under different system-wide power-constraints. For such power-constrained systems, improving per-job flops-per-watt may not be sufficient to improve the total HPC productivity as more number of scientific applications with different compute intensities are migrating to the HPC systems. To measure HPC productivity for such applications, we utilize a monotonically decreasing time-dependent value function, called job-value, with each application. A job-value function represents the value of completing a job for an organization. We begin by exploring the trade-off between two commonly used static power allocation strategies (uniform and greedy) in a power-constrained oversubscribed system. We simulate a large-scale system and demonstrate that, at the tightest power constraint, the greedy allocation can lead to 30% higher productivity compared to the uniform allocation whereas, the uniform allocation can gain up to 6% higher productivity at the relaxed power constraint. We then propose a new dynamic power allocation strategy that utilizes power-performance models derived from offline data. We use these models for reallocating power from running jobs to newly arrived jobs to increase overall system utilization and productivity. In our simulation study, we show that compared to static allocation, the dynamic power allocation policy improves node utilization and job completion rates by 20% and 9%, respectively, at the tightest power constraint. Our dynamic approach consistently earns up to 8% higher productivity compared to the best performing static strategy under different power constraints. Nirmal Kumbhare, Aniruddha Marathe, Ali Akoglu, Salim Hariri, Ghaleb Abdulla |
PDCAT | 3 |
| 2019 | Utility-based resource management in an oversubscribed energy-constrained heterogeneous environment executing parallel applications
Dylan Machovec, Bhavesh Khemka, Nirmal Kumbhare, Sudeep Pasricha, Anthony A. Maciejewski, Howard Jay Siegel, Ali Akoglu, Gregory A. Koenig, Salim Hariri, Cihan Tunc, Michael Wright, Marcia Hilton, Jendra Rambharos, Christopher Blandin, Farah Fargo, Ahmed Louri, Neena Imam |
Parallel Comput. | 7 |
| 2019 | Implementation of scalable bidomain-based 3D cardiac simulations on a graphics processing unit cluster
Ehsan Esmaili, Ali Akoglu, Salim Hariri, Talal Moukabary |
J. Supercomput. | 2 |
| 2018 | Real-Time GPU Based Video Segmentation with Depth InformationabstractIn the context of video segmentation with depth sensor, prior work maps the Metropolis algorithm, a simulated annealing based key routine during segmentation, onto an Nvidia Graphics Processing Unit (GPU) and achieves real-time performance for 320×256 video sequences. However that work utilizes depth information in a very limited manner. This paper presents a new GPU-based method that expands the use of depth information during segmentation and shows the improved segmentation quality over the prior work. In particular, we discuss various ways to restructure the segmentation flow, and evaluate the impact of several design choices on throughput and quality. We introduce a scaling factor for amplifying the interaction strength between two spatially neighboring pixels and increasing the clarity of borderlines. This allows us to reduce the number of required Metropolis iterations by over 50% with the drawback of over-segmentation. We evaluate two design choices to overcome this problem. First, we incorporate depth information into the perceived color difference calculations between two pixels, and show that the interaction strengths between neighboring pixels can be more accurately modeled by incorporating depth information. Second, we pre-process the frames with Bilateral filter instead of Gaussian filter, and show its effectiveness in terms of reducing the difference between similar colors. Both approaches help improve the quality of the segmentation, and the reduction in Metropolis iterations helps improve the throughout from 29 fps to 34 fps for 320×256 video sequences. Nilangshu Bidyanta, Ali Akoglu |
AICCSA | 2 |
| 2018 | Balancing the learning ability and memory demand of a perceptron-based dynamically trainable neural network
Edward Richter, Spencer Valancius, Josiah McClanahan, John Mixter, Ali Akoglu |
J. Supercomput. | 5 |
| 2017 | Multi-Mode Low-Latency Software-Defined Error Correction for Data CentersabstractFlash memories are gaining prominence for utilizing in large scale data centers (DCs) due to their high memory density, low power consumption and heat dissipation, and high access speed characteristics. The rate of degradation for a flash memory is largely affected by the amount and frequency of the erase/write operations, which is a challenge in the DC context that serves dynamically changing workloads. Adaptive Error Correction Code (AECC) schemes have been introduced for changing the error correction algorithm based on the reliability state of the flash. In this study we show that hard decision (bit-flipping) and soft decision decoding (Belief Propagation) class of algorithms for Low Density Parity Check (LDPC) decoders complement each other for utilizing in the flash based DCs in order to meet the dynamically changing reliability level. We propose a new family of ECC to improve the reliability of flash memory. Our Monte-Carlo simulations and Field Programmable Gate Array (FPGA) based hardware implementation analysis show that LDPC decoders are suitable for balancing the throughput, decoding performance and reliability requirements in DCs. Fakhreddine Ghaffari, Ali Akoglu, Bane Vasic, David Declercq |
ICCCN | 2 |
| 2016 | Just In Time Architecture (JITA) for dynamically composable data centersabstractComputer manufacturers, software developers, and service providers spend significant time and resources to ensure that they can optimally support one class of applications. However, they cannot cope with the dynamic and continuous changes in applications and workload types, and consequently their offered services become unstable, fragile, and cannot guarantee the required Quality of Service (QoS). Furthermore, it becomes prohibitively expensive to build data centers (DCs) that are optimized for fixed types of workloads and businesses or to scale up for accommodating new growth in heterogeneous workload demands. Consequently, it is critically important that the architecture of next generation DCs is dynamically customizable to support a wide range of application types or businesses under constrained resources. In this paper, we present some design concepts for our Just In Time Architecture (JITA), a novel data center design to overcome the composable DC challenges by dynamically interconnecting DC components into a virtual DC (VDC) that is optimized to the service level objectives for each class of applications. We present how we can use optical waveguide links to build a passive crossbar that directly interconnects all DC resources at the required throughput and latency. Nirmal Kumbhare, Cihan Tunc, Salim Hariri, Ivan B. Djordjevic, Ali Akoglu, Howard Jay Siegel |
AICCSA | 5 |
| 2016 | Resource efficient real-time processing of Contrast Limited Adaptive Histogram EqualizationabstractContextual Contrast Limited Adaptive Histogram Equalization (C-CLAHE) is an effective method for solving the noise amplification effect of the adaptive histogram equalization (AHE), and enhancing the visibility of local details of an image. Even though C-CLAHE has a smaller memory foot print than CLAHE, complexity of the interpolation process increases the computation demand dramatically. Therefore, FPGA based implementations have been limited to CLAHE only. In this study we introduce three key modifications to the C-CLAHE, and for the first time make it feasible to implement on a resource limited FPGA. We restructure the method so that the histogram redistribution stage is realized with fewer number of iterations. We implement contrast limitation calculations earlier during the histogram generation stage instead of during the histogram redistribution stage, which reduces the block RAM demand. We finally mathematically derive an alternative interpolation calculation used during the remapping stage, which reduces the computation complexity in terms of required multipliers by a factor of 2×, without sacrificing the image quality. These algorithmic modifications allowed us to reduce the block RAM demand by a factor of 12×, logic block demand by a factor of 6.7× compared to the state of the art FPGA based CLAHE implementation, and achieve real time processing of 640 × 480 images at a rate of 354 frames per second. Burak Unal, Ali Akoglu |
FPL | 2 |
| 2014 | Overcoming the Limitations Posed by TCR-beta Repertoire Modeling through a GPU-Based In-Silico DNA Recombination AlgorithmabstractThe DNA recombination process known as V(D)J recombination is the central mechanism for generating diversity among antigen receptors such as T-cell receptors (TCRs). This diversity is crucial for the development of the adaptive immune system. However, modeling of all the α β TCR sequences is encumbered by the enormity of the potential repertoire, which has been predicted to exceed 1015sequences. Prior modeling efforts have, therefore, been limited to extrapolations based on the analysis of minor subsets of the overall TCRbeta repertoire. In this study, we map the recombination process completely onto the graphics processing unit (GPU) hardware architecture using the CUDA programming environment to circumvent prior limitations. For the first time, we present a model of the mouse TCRbeta repertoire to an extent which enabled us to evaluate the Convergent Recombination Hypothesis (CRH) comprehensively at peta-scale level on a single GPU. Gregory M. Striemer, Harsha Krovi, Ali Akoglu, Benjamin Vincent, Ben Hopson, Jeffrey Frelinger, Adam Buntzman |
IPDPS | 3 |
| 2013 | FPGA based single cycle, reconfigurable router for NoC applicationsabstractAn FPGA based, single cycle, low latency router design that can be reconfigured to 1-D and 2-D network on chip architectures is proposed. The design is highly scalable and exploits the features provided by any standard FPGA platform and can be easily ported to an ASIC or any other FPGA platform. Due to the highly interconnect-centric nature of NoCs, the built-in resources of an FPGA in terms of routing channels and on chip logic are ideal and provide a well-utilized platform for the router design. Experimental results prove the proposed design is robust and cost effective. Each router consumes a mere 2.08 μW per bit per hop of power in the worst case while achieving high clock rates of 325 MHz easily on the target FPGA device post design synthesis and emulation. Priyank Gupta, Ali Akoglu, Kathleen L. Melde, Janet Roveda |
ISCAS | 2 |
| 2013 | An Analytical Model for Evaluating Static Power of Homogeneous FPGA ArchitecturesabstractAs capacity of the field-programmable gate arrays (FPGAs) continues to increase, power dissipated in the logic and routing resources has become a critical concern for FPGA architects. Recent studies have shown that static power is fast approaching the dynamic power in submicron devices. In this article, we propose an analytical model for relating homogeneous island-style-based FPGA architecture to static power. Current FPGA power models are tightly coupled with CAD tools. Our CAD-independent model captures the static power for a given FPGA architecture based on estimates of routing and logic resource utilizations from a pre-technology mapped netlist. We observe an average correlation ratio (C-Ratio) of 95% and a minimum absolute percentage error (MAPE) rate of 15% with respect to the experimental results generated by the Versatile Placement Routing (VPR) tool over the MCNC benchmarks. Our model offers application engineers and FPGA architects the capability to evaluate the impact of their design choices on static power without having to go through CAD-intensive investigations. Yoon Kah Leow, Ali Akoglu, Susan Lysecky |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2013 | Integration of Net-Length Factor with Timing- and Routability-Driven Clustering Algorithms
Senthilkumar Thoravi Rajavel, Ali Akoglu |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2012 | Cardiac simulation on multi-GPU platform
Venkata Krishna Nimmagadda, Ali Akoglu, Salim Hariri, Talal Moukabary |
J. Supercomput. | 2 |
| 2011 | MO-pack: many-objective clustering for FPGA CADabstractApplications targeting FPGA integrated systems impose strict energy, channel width and delay constraints. We introduce the first many-objective clustering, MO-Pack, that targets these performance metrics concurrently. Detailed performance comparisons over state of the art clustering strategies targeting energy (P-T-VPack), delay (T-VPack), channel width (iRAC), and timing and routability (T-RPack) show that MO-Pack achieves its goals without increasing the logic area. Senthilkumar Thoravi Rajavel, Ali Akoglu |
DAC | 2 |
| 2011 | An analytical energy model to accelerate FPGA logic architecture investigationabstractThere is a pressing need for exploring innovative reconfigurable architectures with the steady growth in the range of FPGA based applications. However, traditional FPGA architecture design methods require time consuming CAD experimentations to identify the most suitable hardware configuration for the target application. Several analytical models have been recently proposed to predict the relative performance of a given set of architectures. Replacing CAD experiments with these analytical models poses as the solution for reducing the complexity of architecture evaluation process. However, among a large set of existing models, an analytical energy model is missing to supplement the architecture evaluation. We argue that energy can be defined as a function of routed wire length and critical path delay. Therefore, we inherit wire length and critical path delay models to derive an analytical energy model for homogeneous FPGA architectures. We evaluate the impact of variations in logic architecture parameters in terms of LUT size, cluster size and inputs per CLB on the energy performance, and show that our energy model accurately captures the trends observed through CAD experiments. An energy model is robust only if its predictions are in agreement with any CAD flow or benchmark suite. We study the robustness of our energy model by varying the seed selection process of placement, optimization goal of clustering and placement, and the nature of the benchmark suite. In all our experimental evaluations, we observe that the energy model accurately captures the performance trends with a high degree of fidelity. Senthilkumar Thoravi Rajavel, Ali Akoglu |
FPT | 2 |
| 2011 | Parallel Implementation of the Irregular Terrain Model (ITM) for Radio Transmission Loss Prediction Using GPU and Cell BE ProcessorsabstractThe Irregular Terrain Model (ITM), also known as the Longley-Rice model, predicts long-range average transmission loss of a radio signal based on atmospheric and geographic conditions. Due to variable terrain effects and constantly changing atmospheric conditions which can dramatically influence radio wave propagation, there is a pressing need for computational resources capable of running hundreds of thousands of transmission loss calculations per second. Multicore processors, like the NVIDIA Graphics Processing Unit (GPU) and IBM Cell Broadband Engine (BE), offer improved performance over mainstream microprocessors for ITM. We study architectural features of the Tesla C870 GPU and Cell BE and evaluate the effectiveness of architecture-specific optimizations and parallelization strategies for ITM on these platforms. We assess the GPU implementations that utilize both global and shared memories along with fine-grained parallelism. We assess the Cell BE implementations that utilize direct memory access, double buffering, and SIMDization. With these optimization strategies, we achieve less than a second of computation time on each platform which is not feasible with a general purpose processor, and we observe that the GPU delivers better performance than Cell BE in terms of total execution time and performance per watt metrics by a factor of 2.3x and 1.6x, respectively. Yang Song 0007, Ali Akoglu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | Net-length-based routability-driven power-aware clusteringabstractThe state-of-the-art power-aware clustering tool, P-T-VPack, achieves energy reduction by localizing nets with high switching activity at the expense of channel width and area. In this study, we employ predicted individual postplacement net length information during clustering and prioritize longer nets. This approach targets the capacitance factor for energy reduction, and prioritizes longer nets for channel width and area reduction. We first introduce a new clustering strategy, W-T-VPack, which replaces the switching activity in P-T-VPack with a net length factor. We obtain a 9.87% energy reduction over T-VPack (3.78% increase over P-T-VPack), while at the same time completely eliminating P-T-VPack's channel width and area overhead. We then introduce W-P-T-VPack, which combines switching activity and net length factors. W-P-T-VPack achieves 14.26% energy reduction (0.31% increase over P-T-VPack), while further improving channel width by up to 12.87% for different cluster sizes. We investigate the energy performance of routability (channel width)-driven clustering algorithms, and show that W-T-VPack consistently outperforms T-RPack and iRAC by at least 11.23% and 9.07%, respectively. We conclude that net-length-based clustering is an effective method to concurrently target energy and channel width. Lakshmi Easwaran, Ali Akoglu |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2010 | Design and evaluation of a self-healing Kepler for scientific workflowsabstractKepler is a popular open source scientific workflow (SWF) as it simplifies the effort required to construct complex data flow models through a visual interface. As the complexity of the workflow applications that will run on heterogeneous distributed systems increases, fault management becomes a critical design issue for large scale scientific and engineering applications. Due to the long execution times of these applications, it is important that they are fault tolerant; i.e. the workflow application can recover gracefully from faults without the need to restart the application from the beginning. The current implementation of Kepler tool does not support fault tolerance or recovery mechanisms. In this paper, we extend the Kepler capabilities to support fault tolerant scientific workflow (FT-SWF) with a checkpoint mechanism where corrective measures are taken seamlessly in an autonomic manner whenever a fault is detected. To the best of our knowledge, this is the first approach on adding autonomic operations to Kepler. We have evaluated the FT-Kepler on a distributed application used by ecosystem researchers. We evaluated the performance of the workflow with hardware and software based fault scenarios in terms of execution time, recovery time, and the checkpoint mechanism overhead. The experimental evaluations indicate that the checkpoint mechanism adds negligible overhead to the total execution time of the workflow and as the fault rate increases, the number of checkpoints should be increased. Arjun Hary, Ali Akoglu, Youssif B. Al-Nashif, Salim Hariri, Darrel Jenerette |
HPDC | 2 |
| 2009 | Accelerated discovery through integration of Kepler with data turbine for ecosystem researchabstractThere is a need for accelerated discovery cycles (ADCs) for integrating experimental and observational data to capture large-scale dynamic ecosystem complexity, to instantly process massive datasets, to test contrasting mechanistic models and to drive the next set of experiments. The overreaching objective is to enable ADCs by coupling advances in computational models and cyber-systems with the unique experimental infrastructure of Biosphere 2 (B2), a large-scale earth system science facility now under management by the University of Arizona. In the context of ADCs, there is a need for software development environment for modeling complex systems and a middleware for data streaming from the field into the models. Kepler is an open source tool that enables the end user to design scientific workflows in order to manage scientific data and perform complex analysis on the data. Ring buffered network bus (RBNB) data turbine is a middleware system that is used to integrate sensor-based environment observing systems with data processing systems. Currently the integration between Kepler and data turbine is limited to reading from the data turbine only. In ADC, multiple hypotheses are tested with different assimilation models. These models run on a distributed computing environment, therefore capability of simultaneous reads and writes to the data turbine is a necessity. In this paper we show how to integrate Kepler with RBNB data turbine to achieve this capability. We also exploit the open-source features of Kepler system and create customized processing models in order to accelerate and automate the experiments in ecosystems research. We describe in further details our implementation approach to enable future studies on Kepler and data turbine integration. Yaser Jararweh, Arjun Hary, Youssif B. Al-Nashif, Salim Hariri, Ali Akoglu, Darrel Jenerette |
AICCSA | 5 |
| 2009 | Parallel implementation of Irregular Terrain Model on IBM Cell Broadband EngineabstractPrediction of radio coverage, also known as radio ldquohear-abilityrdquo requires the prediction of radio propagation loss. The Irregular Terrain Model (ITM) predicts the median attenuation of a radio signal as a function of distance and the variability of the signal in time and in space. Algorithm can be applied to a large amount of engineering problems to make area predictions for applications such as preliminary estimates for system design, surveillance, and land mobile systems. When the radio transmitters are mobile, the radio coverage changes dynamically, taking on a real-time aspect that requires thousands of calculations per second, which can be achieved through the use of recent advances in multicore processor technology. In this study, we evaluate the performance of ITM on IBM Cell Broadband Engine (BE). We first give a brief introduction to the algorithm of ITM and present both the serial and parallel execution manner of its implementation. Then we exploit how to map out the program on the target processor in detail. We choose message queues on Cell BE which offer the simplest possible expression of the algorithm while being able to fully utilize the hardware resources. Full code segment and a complete set of terrain profiles fit into each processing element without the need for further partitioning. Communications and memory management overhead is minimal and we achieve 90.2% processor utilization with 7.9times speed up compared to serial version. Through our experimental studies, we show that the program is scalable and suits very well for implementing on the CELL BE architecture based on the granularity of computation kernels and memory footprint of the algorithm. Yang Song 0007, Jeffrey A. Rudin, Ali Akoglu |
IPDPS | 3 |
| 2009 | Sequence alignment with GPU: Performance and design challengesabstractIn bioinformatics, alignments are commonly performed in genome and protein sequence analysis for gene identification and evolutionary similarities. There are several approaches for such analysis, each varying in accuracy and computational complexity. Smith-Waterman (SW) is by far the best algorithm for its accuracy in similarity scoring. However, execution time of this algorithm on general purpose processor based systems makes it impractical for use by life scientists. In this paper we take Smith-Waterman as a case study to explore the architectural features of Graphics Processing Units (GPUs) and evaluate the challenges the hardware architecture poses, as well as the software modifications needed to map the program architecture on to the GPU. We achieve a 23x speedup against the serial version of the SW algorithm. We further study the effect of memory organization and the instruction set architecture on GPU performance. For that purpose we analyze another implementation on an Intel Quad Core processor that makes use of Intel's SIMD based SSE2 architecture. We show that if reading blocks of 16 words at a time instead of 4 is allowed, and if 64 KB of shared memory as opposed to 16 KB is available to the programmer, GPU performance enhances significantly making it comparable to the SIMD based implementation. We quantify these observations to illustrate the need for studies on extending the instruction set and memory organization for the GPU. Gregory M. Striemer, Ali Akoglu |
IPDPS | 2 |
| 2008 | Concurrent timing based and routability driven depopulation technique for FPGA packingabstractIn FPGA CAD flow, routability driven algorithms have been introduced to improve feasibility of mapping designs onto the underlying architecture; timing and power driven algorithms have been introduced to meet design specifications. A number of techniques have been proposed to tackle routability, timing or power objectives independently during clustering stage. However, there is minimal work that targets multiple optimization goals. In this paper, we evaluate a clustering technique that targets routability and timing goals simultaneously. We combine the timing-driven T-VPack algorithm with a routability-driven non-uniform depopulation scheme (T-RDPack). Our technique keeps clusters on the critical path fully populated, while depopulating other clusters in the design. This approach has been implemented into the versatile place and route (VPR) toolset. We show that, compared to T-VPack, channel width reductions of 11.5%, 19.1%, 24.7% are achieved while incurring an area overhead of 0.6%, 3.1%, 9.1% respectively with negligible increase in critical path delay, exceeding the performance of T-RPack. Audip Pandit, Lakshmi Easwaran, Ali Akoglu |
FPT | 3 |
| 2008 | A hybrid processing element based reconfigurable architecture for hashing algorithmsabstractGiven the high computation demand for cryptography and hashing algorithms there is a need to develop flexible and high performance architectures. This paper proposes a methodology to derive processing elements as a starting point for the state-of-the-art reconfigurable computing and presents a case-study to show that application- specific reconfigurable computing has performance benefits close to fully-custom designs in addition to the intended reconfigurablity. We use hashing algorithms as a case study to propose a novel application-specific reconfigurable architecture based on a balanced mixture of coarse and fine grained processing elements with a tuned interconnect structure. For that purpose we introduce a methodology to derive hybrid grained processing elements and expose both fine and coarse grain parallelism based on a new common and recurring computation pattern extraction tool. After extracting the recurring patterns between SHA-1 and MD5 algorithms, we derive the unified interconnect architecture tailored to the control data dependencies of both the algorithms. That way the amount of reconfiguration on the proposed architecture when switching between the two algorithms is minimized. The proposed reconfigurable architecture is synthesized using the Synopsys design compiler targeted at TSMC 250 nm libraries. We compare its performance with ASIC technology on SHA-1 and MD5 algorithms. Results show that the proposed architecture which is reconfigurable between the two hashing algorithms has frequency of operation close to ASIC implementation of the individual algorithms for iterative and pipelined versions and results with 35% savings in area. Deepak Sreedharan, Ali Akoglu |
IPDPS | 2 |
| 2008 | A coarse grained and hybrid reconfigurable architecture with flexible NoC router for variable block size motion estimationabstractThis paper proposes a novel application-specific hybrid coarsegrained reconfigurable architecture with a flexible network on chip (NoC) mechanism. Architecture supports variable block size motion estimation (VBSME) with much less resources than ASIC based and coarse grained reconfigurable architectures. The intelligent NoC router supports full search motion estimation algorithm as well as other fast search algorithms like diamond, hexagon, big hexagon and spiral. Our model is a hierarchical hybrid processing element based 2D architecture which supports reuse of reference frame blocks between the processing elements through NoC routers. This reduces the transactions from/to the main memory. Proposed architecture is designed with Verilog-HDL description and synthesized by 90 nm CMOS standard cell library. Results show that our architecture reduces the gate count by 7x compared to its ASIC counterpart that only supports full search method. Moreover, the proposed architecture operates at a frequency comparable to ASIC based implementation to sustain 30 fps. Our approach is based on a simple design which utilizes a high-level of parallelism with an intensive data reuse. Therefore, proposed architecture supports run-time reconfiguration for any block size and for any search pattern depending on the application requirement. Ruchika Verma, Ali Akoglu |
IPDPS | 2 |
| 2007 | Methodology and Toolset for ASIP Design and Development Targeting Cryptography-Based ApplicationsabstractNetwork processors utilizing general-purpose instruction-set architectures (ISA) limit network throughput due to latency incurred from cryptography and hashing applications (AES, DES, MD5, and SHA). This paper presents a methodology using computer-aided design, for development of high-performance application-specific instruction-sets (ASIP) targeting applications saturated in repetitive sequential bitwise operations and data-flow dependencies, thus exposing both fine and coarse grain parallelism through a set of recurring pattern extraction tools. These specific instructions, in conjunction with a minimal set of general-purpose instructions, are then incorporated into a simplistic, single-cycle CPU architecture, for a comparison (in CPU cycles) with common general-purpose instruction latencies (Intel P4). We show that the high-performance instruction set derived, based on the proposed methodology, has the potential for facilitating dramatic improvements in performance (over the software kernel implementation) with substantially increased throughput. Results show that up to 93% of the code is utilized by the instructions derived through the developed toolset resulting with up to 2. 7times cycle time improvement. David Montgomery, Ali Akoglu |
ASAP | 2 |
| 2007 | Wirelength Prediction for FPGAsabstractFPGA CAD tools require wirelength predictions to make informed decisions through clustering, placement and routing stages towards power, area or delay based design goals. Unfortunately, there has been minimal work devoted to estimating individual wirelengths early in the CAD flow. Rent's rule can be used to generate a wirelength distribution but cannot be used to predict lengths of individual wires. Hence, this paper explores "structural metrics" that have been found to possess strong predictive qualities in the ASIC domain. To our knowledge this is a first study in the application of these metrics in the FPGA CAD flow. Results show that the studied metrics capture characteristics of placement optimization carried out by VPR, and hence, are good indicators of post-placement wirelengths. Audip Pandit, Ali Akoglu |
FPL | 2 |
| 2007 | Net Length based Routability Driven PackingabstractFPGA CAD requires routability driven algorithms to improve feasibility of mapping designs onto the underlying architecture. At the clustering stage, routability enhancing parameters such as number of absorbed pins, absorbed nets, criticality, connectivity, and number of net terminals have been identified. We argue that net length is an important parameter not explored in this context. In this paper, we show a structural approach to predict individual net lengths based on the intrinsic shortest path length (ISPL) metric, and use it to reduce track count achieved by previous clustering algorithms. For the set of 20 large MCNC benchmarks, we achieve an average channel width reduction of 3.25 tracks with no area overhead compared to previous approaches. The prediction mechanism and modified clustering algorithm have been incorporated into the VPack/VPR (versatile place and route) toolset. This is a first approach that applies apriori net length predictions within the clustering stage to optimize a design for routability. Audip Pandit, Ali Akoglu |
FPT | 2 |
| 2007 | A Highly Parallel FPGA based IEEE-754 Compliant Double-Precision Binary Floating-Point Multiplication AlgorithmabstractThere is increasing demand for fast floating-point arithmetic support to make field programmable gate arrays (FPGAs) a practical option for scientific applications. We propose a new IEEE-754 compliant double-precision floating-point multiplication algorithm that supports denormal numbers, NaN and exception handling. Solution involves bit-level operations with minimum dependency between partial products through a specialized adder tree structure tailored to make use of modular and parallel nature of FPGAs. We achieve maximum operational frequency of 274MHz for mantissa multiplication and 228MHz for the overall system on Xilinx Virtex-4 platform. Our design carries performance benefits similar to ASIC based algorithms; and routing benefits similar to ripple carry array and carry save multipliers. Proposed approach outperforms algorithm and IP-Core solutions in the academia and Xilinx LogiCORE multiplier when no embedded resources are used. Algorithm allows reaching double-double precision level with much less performance degradation and pipelining demand than IP-Core based approaches. Sandeep K. Venishetti, Ali Akoglu |
FPT | 2 |
| 2007 | A Coarse Grained Reconfigurable Architecture for Variable Block Size Motion EstimationabstractThis paper proposes a novel application-specific coarsegrained reconfigurable architecture with a flexible network on chip (NoC) mechanism. This architecture supports variable block size motion estimation (VBSME) with much less resources than other architectures. The intelligent NoC router supports full search motion estimation algorithm as well as other fast search algorithms like diamond and hexagonal search. Our model is a hierarchical hybrid processing element based 2D architecture which supports reuse of search data between the processing elements with the help of NoC routers. Results show that the area occupied by the proposed architecture is about one-seventh of the area occupied by the state of the art ASIC implementation with comparable operational frequency to sustain 30fps. Ruchika Verma, Ali Akoglu |
FPT | 2 |