VLDB 2026 Research / reviewers in the wild / expert
Volodymyr V. Kindratenko
dblp:70/536
· DBLP profile ↗
34ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0002-9336-4756ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 9 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AI Model Serving on HPC InfrastructureabstractWe introduce two open-source frameworks for deploying and running large language models (LLMs) within the constraints of high-performance computing (HPC) clusters managed by SLURM job scheduler. One framework is designed for high-throughput batch inference of LLM requests using an OpenAI-compatible format, enabling users to submit traditional HPC-style workloads. Complementing this, the other framework allows dynamically allocating HPC resources to provide interactive serving of LLMs accessible via API calls. Both frameworks enable researchers to request HPC resources for model inference on-demand, eliminating setup complexities and allowing users to focus exclusively on their core research. Rohan Marwaha, Qinren Zhou, Kastan Day, Volodymyr V. Kindratenko |
eScience | 4 |
| 2025 | Cost-Aware Federated Learning on the CloudabstractWe introduce FedCostAware, a cost-aware scheduling algorithm designed to optimize synchronous federated learning (FL) on cloud spot instances, which addresses the challenges of training on spot instances and different client budgets by employing intelligent management of the lifecycle of spot instances. This approach minimizes idle resource time and overall expenses. Experiments on real-world medical datasets demonstrate that FedCostAware significantly reduces cloud computing costs compared to conventional spot and on-demand schemes, enhancing the accessibility and affordability of FL. Aditya Sinha, Zilinghan Li, Tingkai Liu, Volodymyr V. Kindratenko, Kibaek Kim, Ravi K. Madduri |
eScience | 4 |
| 2025 | Diamond: Harnessing GPU Resources for Scientific Deep LearningabstractModern research computing cyberinfrastructure, such as ACCESS-CI and NAIRR Pilot, offers GPU resources across geographically distributed clusters to accommodate the increasing needs of scientific deep learning (DL) workloads. Even for high-performance computing (HPC) experts, configuring environments and managing DL workloads across supercomputers remain significant barriers. To address these obstacles, we present Diamond, an open-source platform to simplify and streamline the DL lifecycle on HPC. Diamond provides an intuitive graphical interface that abstracts system-level complexity, enabling users to develop, debug, and deploy DL models with minimal overhead. We identify several challenges in building such a platform, including portability, security, and usability, and propose effective architectural solutions to each. Notably, Diamond enables users to share and reuse DL workload environments across systems and collaborators, reducing redundant setup efforts. Experimental results demonstrate that Diamond reduces the time to first successful deployment by an average of 68%, compared to manual configuration with command lines. The Diamond service is available at https://diamondhpc.ai. Haotian Xie, Rohan Marwaha, Minu Mathew, Song Bian 0002, Gengcong Yang, Minghao Yan, Yadu N. Babuji, Owen Price, Yinzhi Wang, Volodymyr V. Kindratenko, Shivaram Venkataraman, Kyle Chard, Ian T. Foster, Zhao Zhang 0007 |
eScience | 10 |
| 2025 | Introducing 3D Representation for Dense Volume-to-Volume Translation via Score FusionabstractIn volume-to-volume translations in medical images, existing models often struggle to capture the inherent volumetric distribution using 3D voxel-space representations, due to high computational dataset demands. We present Score-Fusion, a novel volumetric translation model that effectively learns 3D representations by ensembling perpendicularly trained 2D diffusion models in score function space. By carefully initializing our model to start with an average of 2D models as in existing models, we reduce 3D training to a fine-tuning process, mitigating computational and data demands. Furthermore, we explicitly design the 3D model’s hierarchical layers to learn ensembles of 2D features, further enhancing efficiency and performance. Moreover, Score-Fusion naturally extends to multi-modality settings by fusing diffusion models conditioned on different inputs for flexible, accurate integration. We demonstrate that 3D representation is essential for better performance in downstream recognition tasks, such as tumor segmentation, where most segmentation models are based on 3D representation. Extensive experiments demonstrate that Score-Fusion achieves superior accuracy and volumetric fidelity in 3D medical image super-resolution and modality translation. Additionally, we extend Score-Fusion to video super-resolution by integrating 2D diffusion models on time-space slices with a spatial-temporal video diffusion backbone, highlighting its potential for general-purpose volume translation and providing broader insight into learning-based approaches for score function fusion. Xiyue Zhu, Dou Hoon Kwark, Ruike Zhu, Kaiwen Hong, Yiqi Tao, Shirui Luo, Yudu Li, Zhi-Pei Liang, Volodymyr V. Kindratenko |
ICML | 9 |
| 2024 | UniNet: Accelerating the Container Network Data Plane in IaaS CloudsabstractKubernetes ($K$8s) is a container orchestration plat-form for cloud-based IaaS environments. While it operates on either bare-metal servers or VMs, users prefer VMs for cost savings and agility reasons despite the added network overhead. This overhead, stemming from dual network tunneling at the VM and container levels, degrades performance. To address this, we present UniNet, a SmartNIC-based solution that offloads container-level network tunneling. We designed UniNet to be compatible with leading Container Network Interfaces (CNIs). This approach involves three key elements: (1) transforming VF-based NICs into a container network gateway, (2) offloading the critical path of the data plane functionalities to SmartNICs for enhanced performance and reduced latency, and (3) instituting an isolated control plane that separates VM- and container-level rule insertions, making it tenant-accessible. UniNet boosts CNI throughput by 7.08 x on average, cuts tail latency by 41.6 %, and reduces CPU usage by up to 5.6 x for the receiver and 4.02 x for the sender, respectively. William Gropp, Hubertus Franke, Bharat Sukhwani, Sameh W. Asaad, Jinjun Xiong, Volodymyr V. Kindratenko, Deming Chen |
CLOUD | 8 |
| 2024 | OS4C: An Open-Source SR-IOV System for SmartNIC-Based Cloud PlatformsabstractSmart network interface cards (SmartNICs) are programmable network cards that enable the flexible offloading of network- and application-level functionality. The last several years have seen a significant rise in research related to smart, programmable NICs. Meanwhile, several open-source FPGA-based NIC and networking projects have emerged. However, these projects lack many key features necessary for strong performance in cloud settings. We identify Single Root In-put/Output Virtualization (SR-IOV) as one of the key missing features in these open-source implementations. SR-IOV enables cloud vendors to grant cloud tenants direct access to hardware resources, dramatically reducing the software overheads of device virtualization. We present OS4C, which extends the popular open-source NIC Corundum [1] with support for SR-IOV. We demonstrate that OS4C can improve virtual machine P99.9 network tail latency by up to 17x, throughput by up to 4x, and CPU effort by up to 3.9x compared to software virtualization. On top of this system, we provide a novel weighted round-robin scheduler that enables tenants and providers to control weight distributions and overhaul the Corundum simulation framework to support multi-tenant tests and performance insights. Marissa Lanz, Bill Dai, Martin Ohmacht, Bharat Sukhwani, Hubertus Franke, Volodymyr V. Kindratenko, Deming Chen |
CLOUD | 8 |
| 2024 | Automated Data Management and Learning-Based Scheduling for Ray-Based Hybrid HPC-Cloud Systems
Tingkai Liu, Huili Tao, Yicheng Lu, Zhongbo Zhu, Marquita Ellis, Sara Kokkila Schumacher, Volodymyr V. Kindratenko |
Euro-Par (1) | 7 |
| 2024 | OS4C: An Open-Source SR-IOV System for SmartNIC-based Cloud PlatformsabstractOS4C is the first open-source 100 Gbps NIC research platform that supports SR-IOV. The preliminary design and corresponding throughput gains are presented. Marissa Lanz, Bill Dai, Martin Ohmacht, Bharat Sukhwani, Hubertus Franke, Volodymyr V. Kindratenko, Deming Chen |
FCCM | 8 |
| 2024 | FedCompass: Efficient Cross-Silo Federated Learning on Heterogeneous Client Devices Using a Computing Power-Aware SchedulerabstractCross-silo federated learning offers a promising solution to collaboratively train robust and generalized AI models without compromising the privacy of local datasets, e.g., healthcare, financial, as well as scientific projects that lack a centralized data facility. Nonetheless, because of the disparity of computing resources among different clients (i.e., device heterogeneity), synchronous federated learning algorithms suffer from degraded efficiency when waiting for straggler clients. Similarly, asynchronous federated learning algorithms experience degradation in the convergence rate and final model accuracy on non-identically and independently distributed (non-IID) heterogeneous datasets due to stale local models and client drift. To address these limitations in cross-silo federated learning with heterogeneous clients and data, we propose FedCompass, an innovative semi-asynchronous federated learning algorithm with a computing power-aware scheduler on the server side, which adaptively assigns varying amounts of training tasks to different clients using the knowledge of the computing power of individual clients. FedCompass ensures that multiple locally trained models from clients are received almost simultaneously as a group for aggregation, effectively reducing the staleness of local models. At the same time, the overall training process remains asynchronous, eliminating prolonged waiting periods from straggler clients. Using diverse non-IID heterogeneous distributed datasets, we demonstrate that FedCompass achieves faster convergence and higher accuracy than other asynchronous algorithms while remaining more efficient than synchronous algorithms when performing federated learning on heterogeneous clients. The source code for FedCompass is available at https://github.com/APPFL/FedCompass. Zilinghan Li, Pranshu Chaturvedi, Shilan He, Volodymyr V. Kindratenko, Eliu A. Huerta, Kibaek Kim, Ravi K. Madduri |
ICLR | 6 |
| 2023 | APPFLx: Providing Privacy-Preserving Cross-Silo Federated Learning as a ServiceabstractCross-silo privacy-preserving federated learning (PPFL) is a powerful tool to collaboratively train robust and generalized machine learning (ML) models without sharing sensitive (e.g., healthcare of financial) local data. To ease and accelerate the adoption of PPFL, we introduce APPFLx, a ready-to-use platform that provides privacy-preserving cross-silo federated learning as a service. APPFLx employs Globus authentication to allow users to easily and securely invite trustworthy collaborators for PPFL, implements several synchronous and asynchronous FL algorithms, streamlines the FL experiment launch process, and enables tracking and visualizing the life cycle of FL experiments, allowing domain experts and ML practitioners to easily orchestrate and evaluate cross-silo FL under one platform. APPFLx is available online at https://appflx.link Zilinghan Li, Shilan He, Pranshu Chaturvedi, Trung-Hieu Hoang, Minseok Ryu, Eliu A. Huerta, Volodymyr V. Kindratenko, Jordan D. Fuhrman, Maryellen L. Giger, Ryan Chard, Kibaek Kim, Ravi K. Madduri |
e-Science | 7 |
| 2023 | YouTubePD: A Multimodal Benchmark for Parkinson's Disease AnalysisabstractThe healthcare and AI communities have witnessed a growing interest in the development of AI-assisted systems for automated diagnosis of Parkinson's Disease (PD), one of the most prevalent neurodegenerative disorders. However, the progress in this area has been significantly impeded by the absence of a unified, publicly available benchmark, which prevents comprehensive evaluation of existing PD analysis methods and the development of advanced models. This work overcomes these challenges by introducing YouTubePD -- the first publicly available multimodal benchmark designed for PD analysis. We crowd-source existing videos featured with PD from YouTube, exploit multimodal information including in-the-wild videos, audio data, and facial landmarks across 200+ subject videos, and provide dense and diverse annotations from clinical expert. Based on our benchmark, we propose three challenging and complementary tasks encompassing both discriminative and generative tasks, along with a comprehensive set of corresponding baselines. Experimental evaluation showcases the potential of modern deep learning and computer vision techniques, in particular the generalizability of the models developed on YouTubePD to real-world clinical settings, while revealing their limitations. We hope our work paves the way for future research in this direction. Andy Zhou, Samuel Li, Pranav Sriram, Jiahua Dong 0002, Ansh Sharma, Yuanyi Zhong, Shirui Luo, Volodymyr V. Kindratenko, George Heintz, Christopher M. Zallek, Yu-Xiong Wang |
NeurIPS | 9 |
| 2023 | Critical Minerals Map Feature Extraction Using Deep LearningabstractCritical minerals play a significant role in various areas such as national security, economic growth, renewable energy development, and infrastructure. The assessment of critical minerals requires examining historical scanned maps. The traditional processes of analyzing these scanned maps are labor-intensive, time-consuming, and prone to errors. In this study, we introduce a deep learning technique to help assess critical minerals by automatically extracting digital features from scanned maps. The polygon features extraction is essential for evaluating the concentration and abundance of critical minerals. The extracted polygon features can be used to update existing geospatial databases, conduct further analysis, and support decision-making processes. The proposed U-Net model takes a 6-channel array as input, where legend feature is concatenated with the map image and serves as a prompt, and the model can generate image segmentation based on arbitrary prompts at test time. Our study shows that the modified U-Net model can effectively extract the mining-related polygon regions based on features listed in legends from historic topographic maps. The model achieves a median F1-score of 0.67. This study has the potential to significantly reduce the time and effort involved in manually digitizing geospatial data from historical topographic maps, thus streamlining the overall assessment process. Shirui Luo, Aaron Saxton, Albert Bode, Priyam Mazumdar, Volodymyr V. Kindratenko |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Sparse Spatio-Temporal Neural Network for Large-Scale ForecastingabstractWe introduce sSTNN, a sparse and parallelized version of a spatio-temporal neural network (STNN) that enables training on much larger datasets. First, we introduce the model architecture and discuss the modifications we made to enable the use of a sparse data structure and multi-GPU parallelization. Then we present empirical results that demonstrate sSTNNs ability to train and inference on a dataset 17 times larger than STNN is capable of. Finally, we discuss the effect of sparsification on runtime and present evidence that sSTNN can achieve upwards of 117× reduction in memory usage compared to STNN. Eamon Bracht, Volodymyr V. Kindratenko, Robert J. Brunner |
IEEE Big Data | 2 |
| 2022 | Weakly Supervised Two-Stage Training Scheme for Deep Video Fight Detection ModelabstractFight detection in videos is an emerging deep learning application with today's prevalence of surveillance systems and streaming media. Previous work has largely relied on action recognition techniques to tackle this problem. In this paper, we propose a simple but effective method that solves the task from a new perspective: we design the fight detection model as a composition of an action-aware feature extractor and an anomaly score generator. Also, considering that collecting frame-level labels for videos is too laborious, we design a weakly supervised two-stage training scheme, where we utilize multiple-instance-learning loss calculated on video-level labels to train the score generator, and adopt the self-training technique to further improve its performance. Extensive experiments on a publicly available large-scale dataset, UBI-Fights, demonstrate the effectiveness of our method, and the performance on the dataset exceeds several previous state-of-the-art approaches. Furthermore, we collect a new dataset, VFD-2000, that specializes in video fight detection, with a larger scale and more scenarios than existing datasets. The implementation of our method and the proposed dataset is available at https://github.com/Hepta-Col/VideoFightDetection. Zhenting Qi, Ruike Zhu, Zheyu Fu, Wenhao Chai, Volodymyr V. Kindratenko |
ICTAI | 5 |
| 2022 | Weakly Supervised Segmentation of Buildings in Digital Elevation ModelsabstractThe lack of quality label data is considered one of the main bottlenecks for training machine and deep learning models. Weakly supervised learning using incomplete, coarse, or inaccurate data is an alternative strategy to overcome the scarcity of training data. We trained a U-Net model for segmenting Buildings’ footprints from a high-resolution digital elevation model, using existing label data from the open-access Microsoft building footprints data set. Comparison using an independent, manually labeled benchmark indicated the success of the weak supervision learning as the quality of the model prediction (IoU: 0.876) surpassed that of the original Microsoft data quality (IoU: 0.672) by approximately 20 percent. Moreover, adding extra channels such as elevation derivatives, slope, aspect, and profile curvatures did not enhance the weak learning process as the model learned directly from the original elevation data. Our results demonstrate the value of using existing data for training deep learning models even if they are noisy and incomplete. Aiman Soliman, Shirui Luo, Rauf Makharov, Volodymyr V. Kindratenko |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Exploring HW/SW Co-Design for Video Analysis on CPU-FPGA Heterogeneous SystemsabstractDeep neural network (DNN)-based video analysis has become one of the most essential and challenging tasks to capture implicit information from video streams. Although DNNs significantly improve the analysis quality, they introduce intensive compute and memory demands and require dedicated hardware for efficient processing. The customized heterogeneous system is one of the promising solutions with general-purpose processors (CPUs) and specialized processors (DNN Accelerators). Among various heterogeneous systems, the combination of CPU and FPGA has been intensively studied for DNN inference with improved latency and energy consumption compared to CPU + GPU schemes and with increased flexibility and reduced time-to-market cost compared to CPU + ASIC designs. However, deploying DNN-based video analysis on CPU + FPGA systems still presents challenges from the tedious RTL programming, the intricate design verification, and the time-consuming design space exploration. To address these challenges, we present a novel framework, called EcoSys, to explore co-design and optimization opportunities on CPU-FPGA heterogeneous systems for accelerating video analysis. Novel technologies include 1) a coherent memory space shared by the host and the customized accelerator to enable efficient task partitioning and online DNN model refinement with reduced data transfer latency; 2) an end-to-end design flow that supports high-level design abstraction and allows rapid development of customized hardware accelerators from Python-based DNN descriptions; 3) a design space exploration (DSE) engine that determines the design space and explores the optimized solutions by considering the targeted heterogeneous system and user-specific constraints; and 4) a complete set of co-optimization solutions, including a layer-based pipeline, a feature map partition scheme, and an efficient memory hierarchical design for the accelerator and multithreading programming for the CPU. In this article, we demonstrate our design framework to accelerate the long-term recurrent convolution network (LRCN), which analyzes the input video and output one semantic caption for each frame. EcoSys can deliver 314.7 and 58.1 frames/s by targeting the LRCN model with AlexNet and VGG-16 backbone, respectively. Compared to the multithreaded CPU and pure FPGA design, EcoSys achieves$20.6\times $and$5.3\times $higher throughput performance. Xiaofan Zhang 0001, Jinjun Xiong, Wen-Mei W. Hwu, Volodymyr V. Kindratenko, Deming Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2012 | Application accelerators in HPC - Editorial introduction
Volodymyr V. Kindratenko, Gregory D. Peterson |
Parallel Comput. | 1 |
| 2012 | Fine-grain parallelism using multi-core, Cell/BE, and GPU Systems
Frederico Pratas, Pedro Trancoso, Leonel Sousa, Alexandros Stamatakis, Guochun Shi, Volodymyr V. Kindratenko |
Parallel Comput. | 6 |
| 2011 | GPU acceleration of an image characterization algorithm for document similarity analysisabstractThis paper aims to provide GPU acceleration of an decision support for selecting software and hardware architecture for content-based document comparison. We evaluate Java, C, CUDA C and OpenCL implementations of an image characterization algorithm used for content-based document comparison on a CPU and NVIDIA and AMD graphics processing units (GPUs). Based on our experimental results, we conclude that the original Java implementation of the image characterization algorithm running on a CPU-based architecture can be accelerated by a factor of 6 if the Java code is re-implemented in C, or by a factor of almost 16 if the Java code is re-implemented in CUDA C and run on NVIDIA GTX 480 GPU hardware. We also provide a power efficiency analysis. Guochun Shi, Volodymyr V. Kindratenko, Rob Kooper, Peter Bajcsy |
AICCSA | 2 |
| 2011 | Design of MILC Lattice QCD Application for GPU ClustersabstractWe present an implementation of the improved staggered quark action lattice QCD computation designed for execution on a GPU cluster. The parallelization strategy is based on dividing the space-time lattice along the time dimension and distributing the sub-lattices among the GPU cluster nodes. We provide a mixed-precision floating-point GPU implementation of the multi-mass conjugate gradient solver. Our single GPU implementation of the conjugate gradient solver achieves a 9x performance improvement over the highly optimized code executed on a state-of-the-art eight-core CPU node. The overall application executes almost six times faster on a GPU-enabled cluster vs. a conventional multi-core cluster. The developed code is currently used for running production QCD calculations with electromagnetic corrections. Guochun Shi, Steven A. Gottlieb, Aaron Torok, Volodymyr V. Kindratenko |
IPDPS | 4 |
| 2011 | Guest Editor's Introduction: Special Issue on High-Performance Computing with AcceleratorsabstractThe 12 papers in this special issue on high-performance computing with accelerators discuss a range of different accelerator architectures and applications. David A. Bader, David R. Kaeli, Volodymyr V. Kindratenko |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2010 | Direct self-consistent field computations on GPU clustersabstractWe present an implementation of one of the direct self-consistent-field (DSCF) calculation techniques, the restricted Hartree-Fock method, on a high-performance computing cluster outfitted with graphics processing units (GPUs) and demonstrate its effectiveness and scalability up to 128 cluster nodes on molecules of as many as 1,732 atoms. We discuss the overall parallel application architecture that relies on message passing interface for distributing workload among GPU cluster nodes and POSIX threads to manage the use of GPUs internal to each node. This approach of combining coarse and fine-grain parallelism on a distributed memory system allows to perform DSCF calculations on molecules that up until now have been unattainable due to the excessive computational requirements. Guochun Shi, Volodymyr V. Kindratenko, Ivan S. Ufimtsev, Todd J. Martínez |
IPDPS | 2 |
| 2009 | GPU clusters for high-performance computingabstractLarge-scale GPU clusters are gaining popularity in the scientific computing community. However, their deployment and production use are associated with a number of new challenges. In this paper, we present our efforts to address some of the challenges with building and running GPU clusters in HPC environments. We touch upon such issues as balanced cluster architecture, resource sharing in a cluster environment, programming models, and applications for GPU clusters. Volodymyr V. Kindratenko, Jeremy Enos, Guochun Shi, Michael T. Showerman, Galen Wesley Arnold, John E. Stone, James C. Phillips, Wen-Mei W. Hwu |
CLUSTER | 1 |
| 2009 | Accelerating Cosmological Data Analysis with FPGAsabstractWe present an implementation of a common cosmological data analysis algorithm in DIME-C for the Nallatech H101 FPGA accelerator board. We apply the Reconfigurable Computing Amenability test to investigate the suitability of FPGAs for accelerating this algorithm. We present a parallelization strategy for the FPGA implementation of the algorithm and discuss its realization in DIME-C. The performance improvements of our FPGA implementation are shown to be in a good agreement with the predictions of the Reconfigurable Computing Amenability test. Volodymyr V. Kindratenko, Robert J. Brunner |
FCCM | 1 |
| 2009 | Phoenix: A Runtime Environment for High Performance Computing on Chip MultiprocessorsabstractExecution of applications on upcoming high-performance computing (HPC) systems introduces a variety of new challenges and amplifies many existing ones. These systems will be composed of a large number of ldquofatrdquo nodes, where each node consists of multiple processors on a chip with symmetric multithreading capabilities, interconnected via high-performance networks. Traditional system software for parallel computing considers these chip multiprocessors (CMPs) as arrays of symmetric multiprocessing cores, when in fact there are fundamental differences among them. Opportunities for optimization on CMPs are lost using this approach. We show that support for fine-grained parallelism coupled with an integrated approach for scheduling of compute and communication tasks is required for efficient execution on this architecture. We propose Phoenix, a runtime system designed specifically for execution on CMP architectures to address the challenges of performance and programmability for upcoming HPC systems. An implementation of message passing interface (MPI) atop Phoenix is presented. Micro-benchmarks and a production MPI application are used to highlight the benefits of our implementation vis-a-vis traditional MPI implementations on CMP architectures. Avneesh Pant, Hassan Jafri, Volodymyr V. Kindratenko |
PDP | 3 |
| 2008 | Implementation of NAMD molecular dynamics non-bonded force-field on the cell broadband engine processorabstractWe present results of porting an important kernel of a production molecular dynamics simulation program, NAMD, to the Cell/B.E. processor. The non-bonded force-field kernel, as implemented in the NAMD SPEC 2006 CPU benchmark, has been implemented. Both single-precision and double-precision floating-point kernel variations are considered, and performance results obtained on the Cell/B.E., as well as several other platforms, are reported. Our results obtained on a 3.2 GHz Cell/B.E. blade show linear speedups when using multiple synergistic processing elements. Guochun Shi, Volodymyr V. Kindratenko |
IPDPS | 2 |
| 2008 | Reconfigurable Systems Summer Institute 2007
Volodymyr V. Kindratenko, Duncan A. Buell |
Parallel Comput. | 1 |
| 2007 | Mitrion-C Application Development on SGI Altix 350/RC100abstractThis paper provides an evaluation of SGIreg RASC^TM RC100 technology from a computational science software developer's perspective. A brute force implementation of a two-point angular correlation function is used as a test case application. The computational kernel of this test case algorithm is ported to the Mitrion-C programming language and compiled, targeting the RC100 hardware. We explore several code optimization techniques and report performance results for different designs. We conclude the paper with an analysis of this system based on our observations while implementing the test case. Overall, the hardware platform and software development tools were found to be satisfactory for accelerating computationally intensive applications, however, several system improvements are desirable. Volodymyr V. Kindratenko, Robert J. Brunner, Adam D. Myers |
FCCM | 1 |
| 2006 | A case study in porting a production scientific supercomputing application to a reconfigurable computerabstractThis case study presents the results of porting a production scientific code, called NAMD, to the SRC-6 high-performance reconfigurable computing platform based on field programmable gate array (FPGA) technology. NAMD is a molecular dynamics code designed to run on large supercomputing systems and used extensively by the computational biophysics community. NAMD's computational kernel is highly optimized to run on conventional von Neumann processors; this presents numerous challenges to its reimplementation on FPGA architecture. This paper presents an overview of the SRC-6 architecture and the NAMD application and then discusses the challenges, solutions, and results of the porting effort. The rationale in choosing the development path taken and the general framework for porting an existing scientific code, such as NAMD, to the SRC-6 platform are presented and discussed in detail. The results and methods presented in this paper are applicable to the large class of problems in scientific computing Volodymyr V. Kindratenko, David Pointer |
FCCM | 1 |
| 2006 | M03 - Reconfigurable supercomputingabstractThe synergistic advances in high-performance computing and reconfigurable computing, based on field programmable gate arrays (FPGAs), has resulted in hybrid parallel systems of microprocessors and FPGAs. Such systems support both fine-grain and coarse-grain parallelism, and can dynamically tune their architecture to fit various applications. Programming these systems can be quite challenging as programming of FPGA devices can involve hardware design. This tutorial will introduce the field of reconfigurable supercomputing and its advances in systems, programming, applications and tools. Reconfigurable system developments at SRC, Cray, SGI, and Star Bridge will be highlighted, and case studies including full application developments will be presented along with the live demonstrations. This tutorial will be the first to show scalability studies for real-life applications over entire HPRC systems. This will reveal the tremendous promise held by this class of architectures in performance, power and cost improvements. Challenges that remain will be also discussed. Tarek A. El-Ghazawi, Duncan A. Buell, Volodymyr V. Kindratenko, Kris Gaj |
SC | 3 |
| 2006 | Reconfigurable supercomputing - Is high-performance reconfigurable computing the next supercomputing paradigm?abstractHigh-Performance Reconfigurable Computers (HPRCs) based on the combination of conventional processors and FPGAs have been gaining attention in the past few years. Their benefits were particularly harnessed in compute-intensive integer applications. However, there has been doubt that the same benefits can be attained for general scientific applications. Fortunately, the trend in reconfigurable chip sizes and diversity of resources may be relieving some of those concerns. Yet, with the hardware reconfigurability, it is feared that domain scientists have to learn how to design hardware if they were to use such machines effectively. In order to address the overarching question, this panel will address the following questions: Can FPGAs deliver order-of-magnitude performance gains in scientific floating-point applications in the foreseeable future? Can programming HPRCs programmability become similar to that of HPCs in its level of difficulty? What are the major developments in the industry or the community that make all this possible? Tarek A. El-Ghazawi, Dave Bennett, Daniel S. Poznanovic, Allan Cantle, Keith D. Underwood, Rob Pennington, Duncan A. Buell, Alan D. George, Volodymyr V. Kindratenko |
SC | 9 |
| 2005 | The Visible Radio: Process Visualization of a Software-Defined RadioabstractIn this case study, a data-oriented approach is used to visualize a complex digital signal processing pipeline. The pipeline implements a frequency modulated (FM) software-defined radio (SDR). SDR is an emerging technology where portions of the radio hardware, such as filtering and modulation, are replaced by software components. We discuss how an SDR implementation is instrumented to illustrate the processes involved in FM transmission and reception. By using audio-encoded images, we illustrate the processes involved in radio, such as how filters are used to reduce noise, the nature of a carrier wave, and how frequency modulation acts on a signal. The visualization approach used in this work is very effective in demonstrating advanced topics in digital signal processing and is a useful tool for experimenting with the software radio design. Alex Betts, Donna J. Cox, David Pointer, Volodymyr V. Kindratenko |
IEEE Visualization | 5 |
| 2003 | IntelliBadgeTM: Towards Providing Location-Aware Value-Added Services at Academic Conferences
Donna J. Cox, Volodymyr V. Kindratenko, David Pointer |
UbiComp | 2 |
| 1996 | Classification of irregularly shaped micro-objects using complex Fourier descriptorsabstractA new method for characterizing complex irregular shapes of natural objects (e.g. plant cells, aerosol particles) is presented. The method employs complex Fourier analysis rather than the traditional forms of the Fourier analysis. Digitized images of objects are processed and the contours are tracked by a classical boundary following technique. A new contour preprocessing technique allows to normalize position, size and orientation of a contour. A new contour resampling technique results in a more precise polygonal approximation of the contour. The resampled contour is represented in a complex plane and its complex Fourier coefficients are computed. The technique is applied to the classification of individual algae cells and their agglomerates employing two different classification methods: hierarchical clustering and neural network. The results demonstrate the applicability of the method for the classification of complex irregular shapes. Volodymyr V. Kindratenko, Piet J. M. Van Espen |
ICPR | 1 |